arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DualG-MRAG:面向多模态检索增强生成的宏观推理与微观匹配解耦框架

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation

Jiacheng Tao, Qingyun Sun, Haonan Yuan, Ziwei Zhang, Jianxin Li

arXiv 2607.28580首次发表:更新:

发表机构

Beihang University; SKLCCSE(北京航空航天大学; SKLCCSE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态RAG在复杂多跳推理中存在的问题,本文提出DualG-MRAG框架,通过解耦宏观推理与微观匹配抑制检索噪声,经实验验证其在证据召回和问答准确率上优于基线。

AI 中文摘要

多模态检索增强生成(MM-RAG)虽已取得良好效果,但在复杂多跳推理任务上仍存在不足。现有方法主要聚焦于独立实例级匹配,往往无法捕捉跨模态及跨文档的显式关系。尽管图增强方法引入了结构建模,但在多模态场景中面临核心难题:融入细粒度视觉特征会导致图快速扩张并产生检索噪声,而粗粒度表示则会丢失关键局部证据。为解决该困境,本文提出DualG-MRAG,这是一个包含宏观推理图与微观匹配图的解耦架构,用于多模态RAG。具体而言,为通过分离全局结构推理与细粒度证据匹配来抑制检索噪声,本文构建用于全局拓扑路由的宏观图和用于精确局部验证的微观图;随后,为实现跨异构证据源的动态相关性传播,本文将检索建模为基于查询的消息传递过程,采用GNN Retriever完成;此外,为向生成模型提供连贯的结构指导,本文引入动态规划解码机制,直接从GNN前向传播中提取显式推理路径,替代孤立文档块的标准输入。大量实验表明,DualG-MRAG在证据召回率和复杂问答准确率上均优于基线方法。

英文摘要

While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex multi-hop reasoning tasks. Existing methods primarily focus on independent instance-level matching, which often fails to capture explicit relationships across modalities and documents. Although Graph-enhanced methods introduce structural modeling, they face a fundamental challenge in multimodal scenarios: incorporating fine-grained visual features leads to rapid graph expansion and retrieval noise, whereas coarse-grained representations cause the discarding of critical local evidence. To address this dilemma, we propose DualG-MRAG, a Dual-tier framework that introduces a decoupled architecture comprising Macro-reasoning and Micro-matching Graphs for Multimodal RAG. Specifically, to suppress retrieval noise by isolating global structural reasoning from fine-grained evidence matching, we construct a Macro Graph for global topological routing and a Micro Graph for precise local verification. Subsequently, to enable dynamic relevance propagation across heterogeneous evidence sources, we formulate retrieval as a query-driven message passing process via a GNN Retriever. Furthermore, to provide the generative model with coherent structural guidance, we introduce a dynamic programming decoding mechanism that extracts explicit reasoning paths directly from the GNN's forward pass, replacing the standard input of isolated document chunks. Extensive experiments demonstrate that DualG-MRAG outperforms baselines in both evidence recall and complex QA accuracy.

CommentsAccepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑