发表机构
Northwestern Polytechnical University; Shanghai Artificial Intelligence Laboratory; The University of Hong Kong; Monash University; The Chinese University of Hong Kong (Shenzhen)(西北工业大学; 上海人工智能实验室; 香港大学; 莫纳什大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有GraphRAG方法将知识图视为静态数据结构的局限,提出EvoGraph-R1自进化框架,把检索建模为MDP,智能体通过多种操作使超图结构进化以支持多跳推理,实验证明该方法在多模态问答基准测试中大幅优于现有基线。
AI 中文摘要
检索增强生成(RAG)已成为将多模态大语言模型(MLLMs)与外部知识相结合的关键范式。近期的GraphRAG方法引入结构化实体关系图来改进检索和推理。然而,它们将知识图视为离线构建并单次查询的静态数据结构,存在局限性。本文提出EvoGraph-R1,一个自进化的GraphRAG框架,将知识图重新概念化为通过智能体交互塑造的动态环境。将检索制定为马尔可夫决策过程(MDP),智能体通过观察图状态执行查询、扩展、细化或终止推理等操作,这些操作重塑超图结构并生成反馈信号以指导后续进化。在多模态VQA和文本QA基准测试上的实验表明,该方法在准确性、覆盖率和可追溯性方面比现有RAG基线有显著改进,确立了自进化知识图作为跨模态的基本范式。
英文摘要
Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Recent GraphRAG methods introduce structured entity-relation graphs to improve retrieval and reasoning. However, they remain limited by treating knowledge graphs as static data structures built offline and queried in a single pass. This static paradigm misaligns with the interactive, iterative nature of knowledge-intensive reasoning, creating three bottlenecks: (i) text-centric fragmentation that impedes cross-modal reasoning, (ii) frozen structures unable to incorporate new evidence or correct errors, and (iii) rigid single-pass retrieval without adaptive refinement. To overcome these limitations, we introduce EvoGraph-R1, a self-evolving GraphRAG framework that reconceptualizes knowledge graphs as dynamic environments shaped through agent interactions. We formulate retrieval as a Markov Decision Process (MDP) where the agent observes the graph state and executes actions to query (GraphRetrieve), expand (WebSearch), refine (GraphEdit), or terminate (Answer) the reasoning. These actions reshape the hypergraph structure and generate feedback signals that guide subsequent evolution. Through this closed loop, the hypergraph evolves by integrating new evidence, correcting errors, and refining structure to support multi-hop reasoning. Experiments on multimodal VQA and text QA benchmarks demonstrate substantial improvements over existing RAG baselines in accuracy, coverage, and traceability, establishing self-evolving knowledge graphs as a fundamental paradigm across modalities.
Comments10 pages main paper, 6 figures. CVPR 2026 accepted paper