arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Doc-REFRAG:重新思考多模态文档检索增强生成

Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation

Ruofan Hu, Shengyang Xu, Minjie Hong, Xiaoda Yang, Sashuai Zhou, Ke Lei, Tao Jin, Zhou Zhao

arXiv 2608.30163首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多模态RAG模型在多图像场景中准确率有限且计算开销大的问题,提出基于大规模数据集DocLongRAG的问题引导框架Doc-REFRAG,其在六个基准上优于11个基线,实现SOTA准确率且推理延迟更低。

AI 中文摘要

现实世界的知识存在于多模态文档中,这使得检索增强生成(RAG)对于准确的问答任务至关重要。然而,现有的多模态RAG模型主要针对单图像或封闭文档场景设计,在现实的多图像场景中准确率有限。此外,处理大量检索到的图像会因无关视觉令牌产生大量计算开销。为解决这些挑战,我们引入DocLongRAG,这是一个包含34.3万对问答的大规模数据集,每对平均关联37.4个检索图像,以反映真实的RAG工作流程。基于该数据集,我们提出Doc-REFRAG,这是一个问题引导的框架,它将视觉令牌压缩为粗块,并通过轻量级基于强化学习(RL)的选择器选择性扩展与问题相关的块。在六个基准上的实验表明,Doc-REFRAG优于11个强大的基线,实现了最先进的准确率,同时推理延迟显著更低。我们的资源可在该https URL获取。

英文摘要

Real-world knowledge resides in multimodal documents, necessitating retrieval-augmented generation (RAG) for accurate question answering. However, existing multimodal RAG models are primarily designed for single-image or closed-document settings and exhibit limited accuracy in realistic multi-image scenarios. Moreover, processing numerous retrieved images incurs substantial computational overhead from irrelevant visual tokens. To address these challenges, we introduce DocLongRAG, a large-scale dataset of 343K question--answer pairs, each associated with an average of 37.4 retrieved images to reflect authentic RAG workflows. Building on this dataset, we propose Doc-REFRAG, a question-guided framework that compresses visual tokens into coarse chunks and selectively expands question-relevant ones via a lightweight RL-based selector. Experiments on six benchmarks show that Doc-REFRAG outperforms eleven strong baselines, achieving state-of-the-art accuracy with significantly lower inference latency. Our resources are available at https://github.com/Collab-Gen/Doc-REFRAG.

CommentsAccepted by EMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑