arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于超图的多模态检索增强生成与增量优化

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao

arXiv 2608.16628首次发表:更新:

发表机构

Hangzhou Dianzi University; Beijing University of Technology; Tsinghua University(杭州电子科技大学; 北京工业大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有多模态检索增强生成系统的二元连接局限与优化冗余问题,提出Hyper-M2RAG框架,通过多模态超图建模与锚点驱动的增量优化机制,在多模态基准数据集上实现优于现有方法的性能。

AI 中文摘要

现代多模态检索增强生成(M-RAG)系统受限于传统简单图的二元连接范式,无法捕获异构实体间复杂的高阶关联,例如视觉图表、分散文本描述与底层数值数据间的N元关系。此外,现有优化策略常依赖详尽的整页重建来对齐跨模态信息,导致计算冗余过高且在长文档处理中引入上下文噪声。本文提出Hyper-M2RAG框架,通过高阶超图表示学习重新定义多模态文档检索:首先将文档结构形式化为多模态超图,利用超边作为统一语义容器封装文本、图像与表格间的多向关联,突破点对点建模的局限;为缓解物理分页导致的语义碎片化,引入锚点驱动的增量优化机制,不进行全局扫描,而是识别跨页锚点节点并利用其一跳邻域上下文重构局部超拓扑,以极小的计算开销有效弥合跨页知识缺口。在多模态基准数据集上的大量评估表明,Hyper-M2RAG在检索精度与生成连贯性上均显著优于现有最优方法,代码可获取于this https URL。

英文摘要

Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling. To mitigate semantic fragmentation caused by physical pagination, we introduce an Anchor-driven Incremental Refinement mechanism. Rather than performing a global sweep, our approach identifies boundary-crossing anchor nodes and reconstructs their local hyper-topology using one-hop neighborhood contexts. This targeted refinement effectively bridges cross-page knowledge gaps with minimal computational footprints. Extensive evaluations on multimodal benchmarking datasets demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods in both retrieval precision and generation coherence. Our code is available at https://github.com/ShenAoChen2001/MMHRAG.

CommentsAccepted to the 34th ACM International Conference on Multimedia (ACM MM 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑