发表机构
Mazelone(马泽隆)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉丰富的长文档问答的多模态检索增强生成,提出VLD-RAG框架,构建多模态索引并采用混合检索策略,经智能体工作流程协调,在相关基准上提升了证据页面检索及问答表现,凸显协调验证与混合检索的重要性。
AI 中文摘要
视觉丰富的文档,如报告、幻灯片和手册,往往将回答问题所需的证据分散在多页中,文本与布局线索、表格、图表和图形混合在一起。这项工作研究了针对此类视觉丰富的长文档进行问答的多模态检索增强生成,其中检索必须选择包含文本和视觉信号的证据页面。我们提出了VLD-RAG,这是一个用于长文档多页证据检索和跨页推理的智能多模态RAG框架。VLD-RAG构建了一个保留页面的多模态索引,存储解析后的文本、页面级元数据和密集视觉表示,并使用一种混合检索策略,将基于关键字的稀疏搜索与密集语义查询相结合,以识别候选源和证据页面。一个验证器引导的智能工作流程协调检索智能体、答案智能体和验证智能体,以扩大证据覆盖范围、检测缺失引用,并在需要时优化检索请求。我们用Top-1和Top-5证据页面准确率评估检索,用广义准确率评估生成,并表明VLD-RAG在包括LongDocURL和MMLongBench-Doc在内的视觉丰富的长文档基准上,提高了证据页面检索和最终任务问答,优于以前基于视觉的检索基线。这些发现突出了在正确答案取决于分散在各页的证据时,协调的智能体验证和多模态混合检索对于可靠基础的至关重要性。
英文摘要
Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal retrieval-augmented generation for question answering over such visually-rich long documents, where retrieval must select evidence pages that include both textual and visual signals. We present VLD-RAG, an agentic multimodal RAG framework for multi-page evidence retrieval and cross-page reasoning over long documents. VLD-RAG builds a page-preserving multimodal index that stores parsed text, page-level metadata, and dense visual representations, and uses a hybrid retrieval strategy that combines keyword-based sparse search with dense semantic queries to identify candidate sources and evidence pages. A verifier-guided agent workflow coordinates a Retrieval Agent, Answer Agent, and Validation Agent to broaden evidence coverage, detect missing citations, and refine retrieval requests when needed. We evaluate retrieval with Top-1 and Top-5 evidence-page accuracy and generation with generalized accuracy, and show that VLD-RAG improves both evidence-page retrieval and end-task question answering on visually-rich long-document benchmarks, including LongDocURL and MMLongBench-Doc, outperforming previous vision-based retrieval baselines. These findings highlight that coordinated agent verification and multimodal hybrid retrieval are crucial for reliable grounding when correct answers depend on evidence scattered across pages.