发表机构
Amazon Inc.(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态视觉问答中模型读取文档不可靠的问题,提出Q-Guide智能体引导感知,在两个数据集上优于基线方法,且增益来自精准感知引导而非复杂控制逻辑。
AI 中文摘要
多模态大语言模型(Multimodal LLMs)能够识别文档,但在可靠读取文档方面仍存在不足:即使文档已处于模型上下文窗口中,小型文本、表格、视觉线索及拓扑元素仍会在直接视觉推理时导致模型出错。多数文档视觉问答(document-VQA)系统将感知过程视为固定步骤:对页面进行一次编码,提出问题后,模型仅依据单次快速传递中提取的信息作答。本文认为文档VQA需要更缓慢、更审慎的感知方式:模型不应基于单次固定编码作答,而应在推理阶段投入额外计算资源,确定后续需关注的内容后再给出答案。我们将该思路融入Q-Guide(一款小型智能体),该智能体读取问题后,会明确自身缺失的证据,调用针对性工具完成证据恢复——在需要文本时读取文本,需要细节时放大查看,或在位置重要时定位对应区域。在DocVQA2026和Manga109数据集上,Q-Guide的性能优于直接提示方法及近期多智能体文档系统:在DocVQA2026上准确率为65.0%,对比基准为40.0%;在Manga109上准确率为32.4%,对比基准为24.4%。该提升在三种Claude主干模型(Opus 4.6、Sonnet 4.6、Opus 4.5)上均成立。我们还发现,准确率随感知预算增加而提升,大部分增益来自2至3轮审慎感知;且增益源于将感知引导至正确位置,而非复杂控制逻辑:添加规划器、路由器或多个协作智能体并无帮助。
英文摘要
Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting in the model's context. Most document-VQA systems treat perception as fixed: they encode the page once, ask the question, and answer from whatever the model happened to extract in that single fast pass. We think document VQA needs slower, more deliberate perception: rather than answering from one fixed encoding, the model should spend a bit of extra compute at inference time working out what to look at next, and only then answer. We build this into \textbf{Q-Guide}, a small agent that reads a question, works out what evidence it is still missing, and calls targeted tool(s) to recover it---reading text where text is needed, zooming in where detail is needed, or grounding a region where position matters. On DocVQA2026 and Manga109, Q-Guide outperforms both direct prompting and recent multi-agent document systems ($65.0\%$ vs.\ $40.0\%$ on DocVQA2026, $32.4\%$ vs.\ $24.4\%$ on Manga109), and the improvement holds across three Claude backbones (Opus 4.6, Sonnet 4.6, and Opus 4.5). We find that accuracy scales with the perception budget---most of the gain appears within two to three deliberate rounds---and that the gain comes from directing perception to the right place, not from complex control logic: adding planners, routers, or multiple collaborating agents does not help.