arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20252cs.CLcs.AI

Lens:为无训练多模态表示学习聚焦正确的语义视角

Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong

AI总结:

Lens提出无训练框架,通过语义视角锚定和上下文短语读出解决语义视角错位,使表示读出任务导向,在MMEB数据集上显著提升性能。

AI中文摘要:

高质量表示对于广泛的下游任务至关重要。专门的嵌入模型被显式优化用于表示学习,但其训练数据在规模和多样性上通常比用于预训练现代大语言模型和多模态大语言模型的庞大数据集更为有限。大规模预训练和指令遵循使自回归模型能够选择相关证据、整合多模态信息,并在不同任务视角下推断语义,为无训练表示学习创造了独特的机会。然而,我们的分析揭示,现有的语义引导方法不能可靠地将提取的状态导向下游任务所需的语义视角。因此,产生的表示往往仍由显著输入内容主导。我们将此问题定义为语义视角错位,并提出Lens,一个无训练框架,使表示读出任务导向。语义视角锚定将任务所需视角与任务特定的读出短语关联,指定后续用于提取的位置的解释角色。上下文短语读出将同一短语置于完整输入之后,并聚合其令牌状态,结合全上下文访问与锚定视角。产生的表示反映任务条件下的证据整合和推断,而非显著内容的通用摘要。无需参数更新、架构修改或重排序,Lens在所有36个MMEB数据集上达到总体Precision@1为63.9,超过最接近的相同骨干无训练嵌入基线10.2个百分点。

英文摘要:

High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity than the massive corpora used to pretrain modern large language models and multimodal large language models. Large-scale pretraining and instruction following enable autoregressive models to select relevant evidence, integrate multimodal information, and infer semantics under different task perspectives, creating a distinctive opportunity for training-free representation learning. However, our analysis reveals that existing semantic-elicitation methods do not reliably orient the extracted states toward the semantic perspective required by the downstream task. Consequently, the resulting representations often remain dominated by salient input content. We characterize this problem as semantic perspective misalignment and propose Lens, a training-free framework that makes representation readout task-directed. Semantic Perspective Anchoring associates the task-required perspective with a task-specific readout phrase, specifying the interpretive role of the positions later used for extraction. Contextualized Phrase Readout places the same phrase after the complete input and aggregates its token states, combining full-context access with the anchored perspective. The resulting representation reflects task-conditioned evidence integration and inference rather than a generic summary of salient content. Without parameter updates, architectural modification, or reranking, Lens achieves an overall Precision@1 of 63.9 across all 36 MMEB datasets, outperforming the closest same-backbone training-free embedding baseline by 10.2 points.

补充信息

↑