发表机构
North Carolina State University; University of North Texas; University of Wyoming(北卡罗来纳州立大学; 北德克萨斯大学; 怀俄明大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TopoGraphRAG-Bench,一个布局锚定的多模态GraphRAG基准,包含2,024个问题,评估系统在文本、图形和表格证据上的拓扑推理能力,发现多模态系统最强但仍需改进布局建模。
AI 中文摘要
真实世界文档将证据分布在复杂页面布局中的文本、表格、图形和标题中。因此,回答此类文档上的复杂问题不仅需要检索相关段落:系统必须恢复连接异构证据单元的证据拓扑。现有的GraphRAG评估主要围绕文本展开,而多模态文档RAG基准评估跨模态检索和生成,但未直接评估预期证据拓扑的恢复。我们引入TOPOGRAPHRAG-BENCH,一个用于GraphRAG中多模态证据推理的布局锚定基准,包含201篇长且视觉丰富的文档上的2,024个问题。问题从文本、图形和表格证据单元自下而上构建,在三种受控拓扑下:单跳检索、桥链推理和多源综合。为确保问题保持其预期结构,我们应用反事实验证以检查捷径抵抗、模态必要性和证据必要性。我们使用检索、生成和拓扑感知推理指标评估纯文本GraphRAG、页面级视觉检索和多模态GraphRAG系统。多模态GraphRAG系统实现了最强的整体性能,但当视觉-文本证据对齐或多单元组合不完整时仍会失败。纯文本GraphRAG在关键依赖基于图形或表格时表现不佳,而页面级视觉检索缺乏拓扑恢复所需的细粒度结构。这些发现促使GraphRAG系统超越文本派生的实体关系图,显式建模文档布局、跨模态证据对齐以及证据单元的推理角色。代码和数据可在以下https URL获取。
英文摘要
Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at https://richardlrc.github.io/TopoGraphRAG-Bench/.
CommentsAccepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)