发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ECHO-$k$,一种任务无关的自监督模态获取方法,利用基础模型内部表示作为代理目标,通过强化学习策略顺序选择信息丰富的模态,在预算约束下提升下游性能。
AI 中文摘要
多模态高维学习的近期进展使得基础模型能够处理异构的大规模数据。然而,在测试时,获取所有特征或模态可能代价高昂且往往冗余。因此,顺序选择信息丰富的模态至关重要,但在下游任务或预测目标未知时颇具挑战。为此,我们提出ECHO-$k$,一种用于模态获取的任务无关且自监督的学习原则:我们利用深度模型内部预训练表示(例如,来自基础模型)作为代理目标,以总结跨模态信息。我们在一个风格化的线性设置中提供了理论保证,该设置激励了用于顺序模态选择的强化学习(RL)策略。在与任务无关和无标签获取基线的对比中,ECHO-$k$在多样化的基础模型后端上持续提升了预算约束下的下游性能。我们的方法为成本感知的测试时部署提供了一条原则性路径,对任何测量昂贵或时间受限且下游任务事先未知的多模态系统都具有意义。
英文摘要
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-$k$, a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model's internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a stylized linear setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across task-agnostic and label-free acquisition baselines, ECHO-$k$ consistently improves budgeted downstream performance across diverse foundation-model backends. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.
CommentsAccepted to NeurIPS 2026
Journal refAdvances in Neural Information Processing Systems, 40 (2026)