AI 中文总结
本研究针对通用多模态检索的需求,提出UMER框架,通过感知对的判别推理联合学习对比式嵌入与判别式排序,在MMEB-V2基准上实现最优性能且支持可调整预算的推理。
AI 中文摘要
通用多模态检索旨在支持各类指令感知的检索任务,既要求语料规模下的高效匹配,又需细粒度的语义推理。近期基于多模态大语言模型(MLLM)的嵌入方法通常从隐藏状态推导表征,而思维链(CoT)推理作为一种有前景的嵌入增强策略,可将中间语义证据编码入表征空间。但现有CoT方法通常对查询与候选进行孤立的项级推理,未提供区分正样本与语义易混淆难负样本的显式证据;此外,对比式嵌入虽能捕捉全局相似性,却难以应对需答案验证、类别判断或细粒度推理的元任务。本文提出UMER,即用于通用多模态检索的统一多模态嵌入与排序框架。UMER以感知对的判别推理替代项级反思,通过比较查询-候选对识别与指令相关的匹配及差异证据;UMER在单个MLLM内联合学习用于高效全局匹配的对比式嵌入,以及用于显式成对相关性判断的判别式排序;互补的互蒸馏策略进一步在嵌入与排序功能间传递可靠的成对偏好。在MMEB-V2基准上,UMER在可比实验设置下实现了最优性能,同时支持可调整预算的推理。
英文摘要
Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, existing CoT methods typically use item-wise reasoning over queries and candidates in isolation, providing no explicit evidence to distinguish a positive from a semantically confusable hard negative. Moreover, contrastive embeddings capture global similarity but struggle with meta-tasks requiring answer verification, category judgment or fine-grained reasoning. In this paper, we propose UMER, a Unified Multimodal Embedding and Ranking framework for universal multimodal retrieval. UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence. UMER jointly learns contrastive embeddings for efficient global matching and discriminative ranking for explicit pairwise relevance judgment within a single MLLM. A complementary mutual distillation strategy further transfers reliable pairwise preferences between the embedding and ranking functions. On the MMEB-V2 benchmark, UMER achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.