arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24861cs.IRcs.AI

HVM-GraphRAG:复杂文档上的整体视图多模态图检索增强生成

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document

发表机构吉林大学 · 香港理工大学
查看机构详情
  • Jilin University(吉林大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xin He, Yili Wang, Wenqi Fan, Qing Li, Qinggang Zhang, Yi Chang, Xin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对复杂文档问答中跨模态证据索引不可靠和图遍历昂贵的问题,提出HVM-GraphRAG框架,用整体视图指导图构建,检索时在概念级图上搜索并通过索引访问证据,实验证明其能提升答案性能和检索效率。

中文摘要 AI 辅助

复杂文档的问答需要模型检索和整合分布在遥远文档区域和模态中的证据。多模态GraphRAG通过用图结构组织文档证据提供了一个有前景的方向。然而,现有方法常受不可靠的跨模态证据索引和昂贵的图遍历困扰。为解决这些问题,我们提出HVM-GraphRAG,一个复杂文档上的整体视图多模态GraphRAG框架。它用整体视图指导图构建,减少噪声和冲突的图更新,在概念级图节点和支持的多模态块之间建立可靠索引。检索时在紧凑的概念级图上搜索,通过索引直接访问支持证据,避免在密集实体级图上的昂贵遍历。获得检索证据后,将块重新组织成特定模态组,使问答模型更好地整合异构证据。在三个数据集上的实验表明,HVM-GraphRAG在大多数评估设置中实现了最佳答案性能,同时显著提高了在线检索效率。

英文摘要

Question answering (QA) over complex documents requires models to retrieve and integrate evidence distributed across distant document regions and modalities. Multimodal GraphRAG provides a promising direction by organizing document evidence with graph structures. However, existing methods often suffer from unreliable cross-modal evidence indexing and expensive graph traversal. To address these issues, we propose HVM-GraphRAG, a holistic-view multimodal GraphRAG framework on complex document. HVM-GraphRAG uses a holistic view to guide graph construction, thereby reducing noisy and conflicting graph updates and building reliable indices between concept-level graph nodes and supporting multimodal chunks. During retrieval, HVM-GraphRAG searches over a compact concept-level graph and directly accesses supporting evidence through the constructed index, avoiding costly traversal over dense entity-level graphs. After obtaining the retrieved evidence, HVM-GraphRAG further reorganizes chunks into modality-specific groups, enabling the answering model to better integrate heterogeneous evidence. Experiments on three datasets show that HVM-GraphRAG achieves the best answer performance in most evaluated settings while substantially improving online retrieval efficiency over representative graph-based baselines.

↑