arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14032cs.IRcs.AI

HAM-RAG:面向结构保真交错生成的层级感知多模态检索增强生成

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

发表机构香港科技大学(广州) · ASCETEX国际有限公司 · MOVENSYS公司
另 3 家 · 查看机构详情
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • ASCETEX INTERNATIONAL LIMITED(ASCETEX国际有限公司)
  • MOVENSYS Inc.(MOVENSYS公司)
  • Schneider Electric(施耐德电气)
  • Technical University of Munich(慕尼黑工业大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu, Wenbin Li, Boo-Ho Yang, Rav Lawana, Ziyue Li, Wei Zeng, Fugee Tsung

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有多模态RAG方法丢失结构化文档层级信息的问题,提出HAM-RAG框架,引入HAM-Bench基准验证,使多模态平均性能提升17.3%,为可靠多模态助手提供了层级感知的基础信号。

中文摘要 AI 辅助

现有多模态检索增强生成(RAG)方法常将结构化文档扁平化处理为孤立的文本和图像单元,削弱了忠实证据选择与放置所需的源文档组织结构及局部文本-图像逻辑。本文提出HAM-RAG,即一种面向结构保真交错生成的层级感知多模态RAG框架。HAM-RAG将文档层级作为检索与生成过程中的基础信号,对文本和视觉证据进行上下文建模,并在提示词中保留源文档的位置信息及局部文本-图像关系。我们进一步推出HAM-Bench基准,涵盖Wukong、Wiki、arXiv和Recipe四类数据,对应游戏攻略、网页、科学论文及分步食谱文档。在多种骨干模型上,HAM-RAG较最强的非层级基准方法,多模态平均性能提升17.3%;在Wukong数据集上,HAM-RAG的Img-CBS指标较最强非层级基准提升24.2%,展现出更优的局部文本-图像对齐能力。主要实验与 ablation 研究共同表明,文档层级是实现可靠图像选择、放置及局部文本-图像对齐的关键基础信号。这些结果凸显了层级感知基础对可靠多模态助手的价值,这类助手生成的答案需忠实于结构化文档的组织结构、流程结构及局部文本-图像证据,如技术手册、维护指南和工业标准操作程序(SOP)。代码可在此URL获取。

英文摘要

Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We propose HAM-RAG, a Hierarchy-Aware Multimodal RAG framework for structure-faithful interleaved generation. HAM-RAG uses document hierarchy as a grounding signal across retrieval and generation, contextualizing textual and visual evidence and preserving source position and local text-image relations in the prompt. We further introduce HAM-Bench, covering Wukong, Wiki, arXiv, and Recipe across game walkthroughs, web pages, scientific papers, and step-wise recipe documents. Across multiple backbones, HAM-RAG improves the main multimodal average by 17.3% over the strongest non-hierarchical baseline. On Wukong, HAM-RAG improves Img-CBS by 24.2% over the strongest non-hierarchical baseline, demonstrating substantially better local text-image alignment. The main experiments and ablation study together demonstrate that document hierarchy is a key grounding signal for faithful image selection, placement, and local text-image alignment. These findings highlight the value of hierarchy-aware grounding for reliable multimodal assistants that generate answers faithful to the source organization, procedural structure, and local text-image evidence of structured documents, such as technical manuals, maintenance guides, and industrial SOPs. The code is available at https://github.com/MCCodeAI/HAM-RAG.git.

↑