发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示文档VLM中检索-读取差距,通过配对协议证明提取文本可显著提升准确率,是有效检索的放大器而非替代品。
AI 中文摘要
检索增强的文档问答假设一旦找到正确的页面,视觉语言模型(VLM)就能读取它。我们证明这一假设常常失败,留下了一个检索-读取差距:证据已找到但未被使用。一个配对协议通过仅使用检索到的页面图像的答案与使用相同图像加上其提取文本的答案进行比较,来隔离这一差距。在FoveDoc-Bench(我们的具有可追溯证据的基准)上,检索几乎找到每个证据页面,但添加CPU-OCR文本将严格准确率提高了13到16个百分点。一个精确的文本层大约使增益翻倍,这一增益出现在来自三个家族的六个VLM中,并且在其中一个家族内,随着规模增大而缩小但未闭合。读者可以读取这一证据但无法找到它:其裁剪恢复了大部分文本增益,而页面上的框则没有。相同的协议识别出两个边界。提取文本对文本证据有帮助,但对图表和图形则中性或有害。其优势随着检索退化而缩小,并且相同形式的不相关文本没有增加可检测的增益。提取文本是有效检索的放大器,而非无效检索的替代品。我们的代码可在https://this https URL获取。
英文摘要
Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page images alone with answers from the same images plus their extracted text. On FoveDoc-Bench, our benchmark with traceable evidence, retrieval finds nearly every evidence page, yet adding CPU-OCR text raises strict accuracy by 13 to 16 points. An exact text layer roughly doubles the gain, which appears across six VLMs from three families and, within one family, narrows with scale without closing. The reader can read this evidence but cannot find it: crops of it recover most of the text gain, boxes around it on the page none. The same protocol identifies two boundaries. Extracted text helps on textual evidence but is neutral or harmful on charts and figures. Its advantage shrinks as retrieval degrades, and unrelated text of the same form adds nothing detectable. Extracted text is an amplifier of retrieval that works, not a substitute for retrieval that does not. Our code is available at https://github.com/atoz03/fovedoc-sup.
Comments5 pages, 5 figures, 3 tables. Submitted to ICASSP 2027