arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理重要内容:面向通用多模态嵌入的检索接地推理

Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

Mingzhou Jiang, Peixi Wu, Hang Cheng, Yunhao Zhou, Biao Yang, Wei Yuan, Yun Li, Fan Yang, Wenwu Ou, Honghui He

arXiv 2609.15296首次发表:更新:

发表机构

Tsinghua Shenzhen International Graduate School, Tsinghua University; School of Artificial Intelligence and Data Science, USTC; Kuaishou Technology; College of Future Information Technology, Fudan University(清华大学深圳国际研究生院; 中国科学技术大学人工智能与数据科学学院; 快手科技; 复旦大学未来信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用多模态嵌入中GRPO信用分配不均和CoT推理延迟高的问题,提出ReWAM框架,通过检索感知自蒸馏和自适应推理,实现最先进检索性能并提升高达5倍推理吞吐量。

AI 中文摘要

通用多模态嵌入(UME)学习跨模态的统一表示,使单一模型能够支持多种检索任务。近期方法在生成嵌入前使用思维链(CoT)推理以更好地理解多模态输入,从而应对复杂检索任务,并通过基于检索奖励的GRPO进一步优化该推理过程。然而,两个局限性阻碍了语料库规模的部署。GRPO为所有CoT标记分配相同的优势,未识别区分正样本与负样本的输入支撑声明或证据。此外,在每次嵌入前生成完整CoT会引入大量延迟,即使部分轨迹已提供足够的检索证据。为解决这些局限性,我们提出“推理重要内容”(ReWAM),一种检索接地推理框架,利用检索反馈指导信用分配和推理计算。具体而言,我们引入检索感知自蒸馏(RASD),从区分正样本与检索到的难负样本的输入支撑证据中构建特权指导。一个在策略自教师利用该指导将轨迹级反馈细化为针对检索相关推理的标记级监督。我们进一步开发检索自适应推理(RAI),使用检索置信度头估计部分CoT的剩余检索效用。它提前停止无成效轨迹,并通过推测解码加速有用延续。在MMEB-V2和MRMR上的大量实验表明,ReWAM实现了最先进的检索性能,同时提供比竞争性显式CoT UME方法高达5倍的推理吞吐量。这些结果弥合了检索质量与推理效率之间的差距,使推理增强的UME可大规模部署。

英文摘要

Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing methods overlook the mismatch bettween candidate-aware retrieval supervision and input-only CoT generation: (1)trajectory-level rewards convey retrieval outcomes without explicitly identifying the input-supported evidence that distinguishes the positive from hard negatives; (2) input-only generation cannot directly assess whether further reasoning improves retrieval, potentially producing redundant CoTs with substantial latency. To bridge this gap, we propose Reason What Matters (ReWAM), a retrieval-grounded framework that aligns candidate-aware supervision with input-only generation. Specifically, we introduce Retrieval-Aware Self-Distillation (RASD), which extracts privileged guidance from input-supported facts and evidence distinguishing the positive from hard negatives. Conditioned on this guidance, an on-policy self-teacher provides token-level feedback to refine credit assignment, directing policy updates toward retrieval-relevant reasoning grounded in the input. We further propose Retrieval-Adaptive Inference (RAI), which learns a retrieval-aware stopping criterion from prefix-level retrieval feedback. It stops redundant reasoning without candidate access and uses speculative decoding to further reduce CoT latency. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. ReWAM thus enables high-quality retrieval through efficient input-only reasoning, making explicit CoT practical for corpus-scale multimodal retrieval. The code will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑