发表机构
Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体RAG系统推理深度不足与状态盲区问题,提出图记忆引导框架GraMRAG,结合视觉-文本桥接推理与拓扑感知策略优化,在复杂多模态长时程推理任务上取得最先进性能。
AI 中文摘要
尽管现有的多智能体检索增强生成(RAG)系统在复杂的多模态推理任务上展现出潜力,但在回答知识密集型问题时,它们在推理深度和记忆结构方面仍存在根本性局限,遭受检索不足和状态盲区(state blindness)的困扰。为应对这些局限,我们提出GraMRAG,一种图记忆引导的多智能体RAG框架,它整合了动态多模态记忆图,以实现稳定的多步多模态推理。我们引入一种视觉-文本桥接推理范式,将多尺度实体裁剪与ReAct风格的视觉工具链统一起来,增强了长时程跨模态推理能力。我们进一步构建了一个多模态记忆图,将智能体推理形式化为动态有向无环图(DAG),显式建模动作-观察依赖关系,以缓解状态盲区并抑制冗余检索。此外,我们提出了拓扑感知策略优化(TAPO),利用图拓扑进行关键路径识别和定向节点剪枝,从而在多步推理轨迹中实现细粒度的信用分配。在具有挑战性的多模态基准上的大量实验表明,我们的方法持续优于现有基线,并在复杂长时程推理任务上达到了最先进的性能。
英文摘要
Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they remain fundamentally limited in reasoning depth and memory structure, suffering from inadequate retrieval and state blindness when answering knowledge-intensive questions. To address these limitations, we propose GraMRAG, a graph memory-guided multi-agent RAG framework that integrates a dynamic multimodal memory graph to enable stable, multi-step multimodal reasoning. We introduce a vision-text bridged reasoning paradigm that unifies multi-scale entity cropping with a ReAct-style visual toolchain, enhancing the long-horizon cross-modal reasoning capability. We further construct a multimodal memory graph that formalizes agent reasoning as a dynamic directed acyclic graph (DAG), explicitly modeling action-observation dependencies to mitigate state blindness and suppress redundant retrieval. Moreover, we propose Topology-Aware Policy Optimization (TAPO) that leverages graph topology for critical path identification and targeted node pruning, enabling fine-grained credit assignment across multi-step reasoning trajectories. Extensive experiments on challenging multimodal benchmarks demonstrate that our approach consistently outperforms existing baselines and achieves state-of-the-art performance on complex long-horizon reasoning tasks.