发表机构
Coventry University(考文垂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究文档问答中多模态和图增强RAG的作用,提出用多模态图-RAG架构,通过独立检索并融合文本、图和视觉证据源来评估效果。经实验发现其价值取决于多种因素,如检索设计等,还揭示了一些相关问题,如图像检索及生成器效率对结果的影响。
AI 中文摘要
检索增强生成(RAG)系统通常基于文档提取的文本运行,可能会丢失图表、布局及跨段落关系中包含的信息。我们提出了一种可解释的多模态图-RAG架构,用从大型语言模型提取的主语-关系-宾语三元组以及基于CLIP的图表检索来增强纯文本基线。这三个证据源独立检索,仅在生成时融合,以便分别评估图证据、视觉证据和生成器选择的效果。我们使用两个封闭权重和两个开放权重的多模态生成器,对1000个PubLayNet页面上的单段落、多跳和图表问题进行了对照四向消融实验。还比较了匹配的可回答标题和仅像素图表问题集,以区分标题恢复和真正的视觉问答。知识图谱增强在该语料库的生成器或问题类型中未提供可靠的准确性提升。在仅像素问题上,纯文本系统准确率为零,多模态系统为0.057 - 0.114,受图像检索(Recall@3 = 0.371)和生成器解释密集科学图表能力的限制。标题衍生问题高估了纯文本视觉问答能力。处理相同图像时,生成器的输入令牌变化达11倍,表明图像令牌化可主导部署成本。这些发现表明多模态和图增强的价值取决于检索设计语料库结构、基准构建和生成器效率。
英文摘要
Graph and multimodal extensions to retrieval-augmented generation (RAG) are often evaluated end to end, making it difficult to isolate whether gains arise from retrieval, prompt-side context, visual access, generator capability, or benchmark construction. We present a stage- and evidence-controlled evaluation across five RAG configurations, four multimodal generators, and three document corpora. The same LLM-extracted knowledge graph is used either after retrieval as provenance-constrained triple injection (+KG) or during retrieval as entity-bridged passage expansion (+KGret). Prompt-side graph injection yields no consistent accuracy improvement and generally reduces faithfulness. In contrast, +KGret increases gold-evidence completeness from 0.22 to 0.46 on HotpotQA bridge questions and from 0.50 to 0.72 on SPIQA cross-paper questions, improving accuracy for every generator on both evidence-deficient sets while having little effect on retrieval-complete controls. For visual question answering, matched caption-answerable and verified pixel-only protocols show that apparent multimodal gains are sensitive to textual leakage. Programmatic checks reveal answer recoverability from captions, corpus text, and model responses generated without complete gold evidence. Accuracy on incomplete-evidence questions reaches 0.35--0.71 on widely disseminated corpora, compared with 0 on PubLayNet, indicating that raw accuracy can overstate retrieval-attributable performance. These results show that graph augmentation is most effective when it changes retrieval under evidence deficits, while multimodal evaluation requires explicit verification that answers are unavailable through text.
Comments14 pages, 6 figures