arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检索前先思考:多模态检索增强生成的智能体规划

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang

arXiv 2607.22643首次发表:更新:

发表机构

Bosch AI Research Center; University of Notre Dame; Pennsylvania State University(博世人工智能研究中心; 圣母大学; 宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态检索增强生成问题,提出MM-R2框架,通过建模检索内容和位置在检索前推理,构建意图基础检索状态,在结构化知识图谱检索,还构建相关数据集并采用两阶段后训练策略,实验证明其在答案准确性等方面表现出色。

AI 中文摘要

多模态检索增强生成(mRAG)旨在利用外部知识回答图像-文本查询,但现有多数系统仍在平坦证据空间上直接从原始多模态输入进行检索。这种设计常面临两个关键挑战:检索目标未充分明确,因问题意图必须基于正确视觉指称;搜索空间结构松散,迫使语义不同的证据在单个全局排序步骤中竞争。我们提出MM-R2,一种多模态智能体检索框架,通过明确建模检索内容和检索位置在检索前进行推理。MM-R2首先从图像-问题对构建意图基础检索状态,捕捉信息需求、基础指称和检索约束。然后在结构化知识图谱上执行检索,智能体在其中选择相关检索单元并在其中发出基础查询。为实现此能力,我们构建了MM-R2-Traj,一个多步检索过程的大规模轨迹数据集,并采用带有监督微调的两阶段后训练策略和GRPO。在Infoseek和百科全书式VQA数据集上的实验表明,MM-R2在答案准确性上显著优于强大基线,同时还产生更具可解释性和可验证性的检索轨迹。

英文摘要

Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles with two key challenges: the retrieval target is under-specified because the question intent must be grounded to the correct visual referent, and the search space is weakly structured, forcing semantically distinct evidence to compete in a single global ranking step. We propose MM-R2, a multimodal agentic retrieval framework that reasons before retrieval by explicitly modeling both what to retrieve and where to search. MM-R2 first constructs an intent-grounded retrieval state from the image-question pair, capturing the information need, grounded referent, and retrieval constraints. It then performs retrieval over a structured KnowledgeMap, where the agent selects relevant retrieval units before issuing grounded queries within them. To enable this capability, we build MM-R2-Traj, a large-scale trajectory dataset of multi-step retrieval processes, and adopt a two-stage post-training strategy with supervised fine-tuning and GRPO. Experiments on Infoseek and Encyclopedic VQA datasets show that MM-R2 substantially outperforms strong baselines on answer accuracy while also yielding more interpretable and verifiable retrieval trajectories.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑