从失败中学习:基于难负例的以检索为中心的思维链用于统一多模态检索
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
中文总结 AI 辅助
该研究提出UniME-R1框架,通过基于难负例的以检索为中心的思维链(RC-CoT)学习从检索失败中优化,在多模态检索基准上提升了性能。
中文摘要 AI 辅助
统一多模态检索旨在识别满足通过异构输入表达的复杂用户意图的候选对象。尽管基于大视觉语言模型(LVLM)的检索器高效且可扩展,但直接编码原始多模态输入通常会丢失细粒度的区分线索,导致语义相似的候选对象之间产生混淆。最近的方法通过生成思维链(CoT)理由来丰富查询表示,以缓解这一限制。然而,这种推理通常仅从查询本身得出:它解释了查询描述的内容,但未说明检索器误解的内容。我们认为,有效的检索推理应基于检索反馈。基于这一见解,我们引入UniME-R1,这是一个嵌入器-顾问框架,用于学习对初始检索到的候选对象进行推理并生成以检索为中心的思维链(RC-CoT)。顾问会单独分析候选对象,以识别嵌入器混淆的区分线索。如果目标出现在初始top-k集合中,UniME-R1会直接对候选对象进行重新排序;否则,它会生成RC-CoT以细化检索方向,并使用双模式嵌入器对整个语料库进行重新检索。为了训练该框架,我们挖掘难负例以模拟现实的检索失败,联合优化直接检索和RC-CoT增强型检索,并通过监督学习和面向检索的强化学习使顾问与检索结果对齐。在MMEB-V2和一系列多样化的通用多模态检索基准上进行的大量实验表明,UniME-R1相较于强大的基线持续提升了检索性能。
英文摘要
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.