发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多模态RAG模型在多图像场景中准确率有限且计算开销大的问题,提出基于大规模数据集DocLongRAG的问题引导框架Doc-REFRAG,其在六个基准上优于11个基线,实现SOTA准确率且推理延迟更低。
AI 中文摘要
现实世界的知识存在于多模态文档中,这使得检索增强生成(RAG)对于准确的问答任务至关重要。然而,现有的多模态RAG模型主要针对单图像或封闭文档场景设计,在现实的多图像场景中准确率有限。此外,处理大量检索到的图像会因无关视觉令牌产生大量计算开销。为解决这些挑战,我们引入DocLongRAG,这是一个包含34.3万对问答的大规模数据集,每对平均关联37.4个检索图像,以反映真实的RAG工作流程。基于该数据集,我们提出Doc-REFRAG,这是一个问题引导的框架,它将视觉令牌压缩为粗块,并通过轻量级基于强化学习(RL)的选择器选择性扩展与问题相关的块。在六个基准上的实验表明,Doc-REFRAG优于11个强大的基线,实现了最先进的准确率,同时推理延迟显著更低。我们的资源可在该https URL获取。
英文摘要
Real-world knowledge resides in multimodal documents, necessitating retrieval-augmented generation (RAG) for accurate question answering. However, existing multimodal RAG models are primarily designed for single-image or closed-document settings and exhibit limited accuracy in realistic multi-image scenarios. Moreover, processing numerous retrieved images incurs substantial computational overhead from irrelevant visual tokens. To address these challenges, we introduce DocLongRAG, a large-scale dataset of 343K question--answer pairs, each associated with an average of 37.4 retrieved images to reflect authentic RAG workflows. Building on this dataset, we propose Doc-REFRAG, a question-guided framework that compresses visual tokens into coarse chunks and selectively expands question-relevant ones via a lightweight RL-based selector. Experiments on six benchmarks show that Doc-REFRAG outperforms eleven strong baselines, achieving state-of-the-art accuracy with significantly lower inference latency. Our resources are available at https://github.com/Collab-Gen/Doc-REFRAG.
CommentsAccepted by EMNLP 2026 Main