HAM-RAG:面向结构保真交错生成的层级感知多模态检索增强生成
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
另 3 家 · 查看机构详情
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- ASCETEX INTERNATIONAL LIMITED(ASCETEX国际有限公司)
- MOVENSYS Inc.(MOVENSYS公司)
- Schneider Electric(施耐德电气)
- Technical University of Munich(慕尼黑工业大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文针对现有多模态RAG方法丢失结构化文档层级信息的问题,提出HAM-RAG框架,引入HAM-Bench基准验证,使多模态平均性能提升17.3%,为可靠多模态助手提供了层级感知的基础信号。
中文摘要 AI 辅助
现有多模态检索增强生成(RAG)方法常将结构化文档扁平化处理为孤立的文本和图像单元,削弱了忠实证据选择与放置所需的源文档组织结构及局部文本-图像逻辑。本文提出HAM-RAG,即一种面向结构保真交错生成的层级感知多模态RAG框架。HAM-RAG将文档层级作为检索与生成过程中的基础信号,对文本和视觉证据进行上下文建模,并在提示词中保留源文档的位置信息及局部文本-图像关系。我们进一步推出HAM-Bench基准,涵盖Wukong、Wiki、arXiv和Recipe四类数据,对应游戏攻略、网页、科学论文及分步食谱文档。在多种骨干模型上,HAM-RAG较最强的非层级基准方法,多模态平均性能提升17.3%;在Wukong数据集上,HAM-RAG的Img-CBS指标较最强非层级基准提升24.2%,展现出更优的局部文本-图像对齐能力。主要实验与 ablation 研究共同表明,文档层级是实现可靠图像选择、放置及局部文本-图像对齐的关键基础信号。这些结果凸显了层级感知基础对可靠多模态助手的价值,这类助手生成的答案需忠实于结构化文档的组织结构、流程结构及局部文本-图像证据,如技术手册、维护指南和工业标准操作程序(SOP)。代码可在此URL获取。
英文摘要
Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We propose HAM-RAG, a Hierarchy-Aware Multimodal RAG framework for structure-faithful interleaved generation. HAM-RAG uses document hierarchy as a grounding signal across retrieval and generation, contextualizing textual and visual evidence and preserving source position and local text-image relations in the prompt. We further introduce HAM-Bench, covering Wukong, Wiki, arXiv, and Recipe across game walkthroughs, web pages, scientific papers, and step-wise recipe documents. Across multiple backbones, HAM-RAG improves the main multimodal average by 17.3% over the strongest non-hierarchical baseline. On Wukong, HAM-RAG improves Img-CBS by 24.2% over the strongest non-hierarchical baseline, demonstrating substantially better local text-image alignment. The main experiments and ablation study together demonstrate that document hierarchy is a key grounding signal for faithful image selection, placement, and local text-image alignment. These findings highlight the value of hierarchy-aware grounding for reliable multimodal assistants that generate answers faithful to the source organization, procedural structure, and local text-image evidence of structured documents, such as technical manuals, maintenance guides, and industrial SOPs. The code is available at https://github.com/MCCodeAI/HAM-RAG.git.