AI 中文总结
本研究实证比较了长文档QA中VLM流水线的输入方式、检索器和智能体/静态选择,发现检索模态比检索器更重要,智能体仅在大型VLM上有效,且不同流水线互补,oracle组合可提升约13个百分点。
AI 中文摘要
视觉语言模型(VLMs)越来越多地被用于长文档处理,其中输入将文本与图表、表格、图形和复杂布局相结合。部署它们意味着要选择如何将文档输入模型,当仅发送部分页面时使用哪个检索器,以及是以智能体方式还是静态流水线方式运行模型。我们在两个长文档问答基准上,使用前沿API和开放权重VLM研究了这些选择。首先,在MMLongBench-Doc上,我们配备页面、表格、图形和搜索调用的六工具智能体仅在回答VLM足够大时才获得回报:使用Qwen3.5-4B和9B时,它落后于静态页面输入;使用Qwen3.5-27B时,它持平;使用Sonnet 4.5时,它领先。在LongDocURL上,它在每个阅读器上都与静态输入持平或领先。它相对于最强静态流水线的领先优势在MMLongBench-Doc上使用前沿阅读器时最为明显,而在LongDocURL上则缩小至噪声范围内。其次,检索模态比具体检索器更重要:最强的图像检索器领先于最强的文本流水线,而在文本方面,单个现成的交叉编码器重排序基本上匹配了更重的多阶段LLM流水线。Top-k图像检索也是我们配对的每个阅读器中token效率最高的输入,大约仅为发送每个页面token的七分之一到四分之一。第三,贯穿这三个选择,我们最强的三个流水线在不同问题上取得成功,一个为每个问题选择最佳流水线的oracle比最佳单一流水线大约高出十三个百分点,尽管证据类型路由几乎无法恢复其中的任何一部分。
英文摘要
Vision-Language Models (VLMs) are increasingly used for long-document processing, where the inputs combine text with charts, tables, figures, and complex layouts. Deploying them means choosing how to feed the document to the model, which retriever to use when only a subset of pages is sent, and whether to run the model agentically or as a static pipeline. We study these choices on two long-document QA benchmarks with both frontier API and open-weight VLMs. First, on MMLongBench-Doc our six-tool agent with page, table, figure, and search calls pays off only once the answering VLM is large enough: with Qwen3.5-4B and 9B it trails static page input, with Qwen3.5-27B it draws level, and with Sonnet 4.5 it leads. On LongDocURL it is level with or ahead of static input at every reader. Its lead over the strongest static pipeline is clearest with the frontier reader on MMLongBench-Doc and narrows to within noise on LongDocURL. Second, retrieval modality matters more than the specific retriever: the strongest image retriever leads the strongest text pipeline, and on the text side a single off-the-shelf cross-encoder rerank essentially matches a much heavier multi-stage LLM pipeline. Top-k image retrieval is also the most token-efficient input at every reader we paired it with, at roughly a seventh to a quarter of the tokens of sending every page. Third, cutting across all three choices, three of our strongest pipelines succeed on different questions, and an oracle that picks the best pipeline per question gains roughly thirteen points over the best single pipeline, though evidence-type routing recovers almost none of it.
Comments22 pages. EMNLP 2026 Industry Track