发表机构
ServiceNow Canada(ServiceNow加拿大公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估多模态文档问答中图像与文本输入表示,发现两者互补,并利用轻量级TF-IDF路由器结合两者优势,提升准确率并降低延迟。
AI 中文摘要
每个文档问答系统都始于一个很少被单独研究的选择:是向模型输入页面图像、提取的文本,还是两者兼用。我们将这一选择独立出来,在四个商业模型端点、两个语料库和两种上下文模式(黄金证据页和完整文档)下,保持提示、评判器和评分流程固定不变。在符合图像预算的文档上,页面图像在两个语料库的每个文档长度上都以准确率领先,但这种优势伴随着不断增长的延迟和成本代价:随着文档变长,文本延迟基本保持平稳,而图像延迟则稳步上升。文本和图像也在不同的问题上失败,在报告的单元格中,恰好只有一种表示正确的项目占19%至25%,因此两者都不能完全替代对方。利用这种互补性,一个仅读取问题文本的轻量级TF-IDF路由器,在文档不相交的保留测试集上,相比始终使用文本的方法获得了2.6个百分点的提升,同时相比始终使用视觉的方法将中位延迟降低了30%。
英文摘要
Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On documents that fit the image budget, page images lead on accuracy at every document length on both corpora, but this advantage carries a growing latency and cost premium: text latency stays roughly flat as documents lengthen while image latency rises steadily. Text and images also fail on different questions, with exactly one representation correct on 19--25% of items across the reported cells, so neither subsumes the other. Exploiting this complementarity, a lightweight TF-IDF router that reads only the question text gains 2.6 points over always-text while cutting median latency 30% relative to always-vision, on a document-disjoint held-out split.
CommentsAccepted at the Context Beyond the Window (CBW) Workshop at COLM 2026 and the DocInsights Workshop at EMNLP 2026