发表机构
University of Oslo; Bosch Center for AI; University of Stuttgart; Hong Kong University of Science and Technology (Guangzhou); Tsinghua University(奥斯陆大学; 博世人工智能中心; 斯图加特大学; 香港科技大学(广州); 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长文档VQA智能体的记忆质量盲区,提出AWM方法,通过融入仅记忆可回答性的GRPO奖励机制,提升了最终答案准确率并降低记忆缺失正确的比例。
AI 中文摘要
长文档视觉问答(VQA)日益依赖VLM智能体,这类智能体需检索候选页面、检查页面图像、将发现内容写入工作记忆并合成答案。工作记忆应在页面检查过程中承载支持答案的证据,以支撑后续基于事实的回答,但现有评估主要仅检查最终答案的正确性和证据页面的访问情况,这形成了记忆质量的盲区:智能体可能访问到正确页面并给出正确答案,但其留下的记忆过于通用或不完整,一旦页面上下文被移除便无法支撑回答。我们提出“仅记忆可回答性”这一诊断指标,用于判断是否仅通过问题和终端工作记忆即可回答问题。基于该诊断指标,可回答工作记忆(AWM)将终端工作记忆视为可回答的证据制品,而AWM-GRPO则将该信号融入GRPO奖励机制,同时保留最终答案的优先级。在GRPO机制下,该奖励会为那些最终答案正确且终端工作记忆仍可回答的轨迹分配更高的优势值。在MMLongBench-Doc数据集上,即使提供了黄金证据页面,仍有42.5%的正确答案无法仅通过终端工作记忆单独回答。AWM-GRPO在MMLongBench-Doc和LongDocURL数据集上,较RAG基线分别将最终答案准确率提升8.1和11.9个百分点,且较仅答案GRPO将记忆缺失正确的比例降低2.7个百分点。
英文摘要
Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-supporting evidence across page inspections for later grounded answering, yet existing evaluation mainly checks final-answer correctness and evidence-page access. This creates a memory-quality blind spot: an agent may reach the right page and answer correctly while leaving behind memory too generic or incomplete to support answering once page context is removed. We introduce \emph{memory-only answerability}, a diagnostic that asks whether a reader can answer from the question and terminal working memory alone. Building on this diagnostic, \emph{Answerable Working Memory} (AWM) treats terminal working memory as an answerable evidence artifact, and AWM-GRPO incorporates this signal into the GRPO reward while preserving final-answer priority. Under GRPO, this reward assigns higher advantages to answer-correct trajectories whose terminal working memory remains answerable. On \textsc{MMLongBench-Doc}, even when gold evidence pages are provided, 42.5\% of correct answers still cannot be answered from terminal working memory alone. AWM-GRPO improves final-answer accuracy over the RAG baseline by 8.1 and 11.9 points on \textsc{MMLongBench-Doc} and \textsc{LongDocURL} and reduces the memory-missing-correct rate by 2.7 points over answer-only GRPO.
CommentsEMNLP 2026 Findings. 16 pages, 4 figures, 9 tables