arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33659cs.CVcs.IR

学习具有证据对齐读出层的多模态嵌入

Learning Multimodal Embeddings with Evidence-Aligned Readout

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Tencent Yuanbao(腾讯元宝)
  • Tsinghua University(清华大学)
  • Tencent(腾讯)
  • The University of Hong Kong(香港大学)
  • University of Tsukuba(筑波大学)

机构由 AI 辅助整理,请以论文原文为准。

Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Jiachuan Wang, Yongqi Zhang

中文总结 AI 辅助

本文提出EviAlign,通过将语义证据组织为五个单元并在边界读取状态,与对比检索目标联合训练,在12个MMEB任务上以500K训练对达到76.9的平均Recall@1,证明证据组织与读出位置的协同设计能提升多模态检索嵌入性能。

中文摘要 AI 辅助

多模态大语言模型能够通过生成暴露与任务相关的证据,但产生有用的证据本身并不能决定它如何进入检索嵌入。我们研究证据的语义组织是否也能指定表示被读取的位置。为了解决这个问题,我们引入了EviAlign,它在共享的多模态大语言模型中将语义证据生成与边界读出层耦合。它将证据组织成五个语义单元,在每个单元边界读取上下文状态,并将这些状态聚合成一个归一化的嵌入。生成和对比检索目标共同训练这个共享结构。使用相同的尾部读出层,语义证据和自由形式的思维链(CoT)产生几乎相同的检索性能,这表明仅凭证据组织并不能解释全部增益。一项受控的2×3研究比较了三种读出策略下一致和打乱的证据组织,使用具有匹配证据跨度的训练目标。使用五个读出状态和相同的平均池化,一致语义组织的优势从基于长度的训练位置的0.65分增长到证据边界的2.39分,产生了1.74分的协同设计交互。在12个MMEB检索任务中,EviAlign在500K训练对下实现了76.9的平均Recall@1,同时保留了单向量索引和评分。

英文摘要

Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled $2\times3$ study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.

↑