AI 中文总结
本文提出DocTrace分层框架,将长文档视觉问答建模为显式证据图推理问题,经两阶段训练优化后,在三个基准上优于现有模型,且推理可追溯。
AI 中文摘要
长文档视觉问答(LongDocVQA)要求多模态大语言模型(MLLM)定位、整合并推理跨多页面分布的异构文档元素。现有方法包括端到端MLLM、检索增强生成(RAG)流水线及文档智能体,往往缺乏显式机制表示和验证推理过程中接地证据的逐步组合方式,限制了答案准确性和可追溯性。本文将LongDocVQA建模为显式证据图推理问题,而非隐式答案预测,为此提出DocTrace,这是一个分层框架,逐步执行证据定位、结构化文档解析和证据图推理,以实现显式证据来源。为有效学习这些能力,我们开发了两阶段训练框架:联合监督微调(SFT)首先初始化证据定位和图推理能力,随后采用带有专用奖励的任务特定组相对策略优化(GRPO)进一步优化这些能力。在MMLongBench-Doc、LongDocURL和SlideVQA上的大量实验表明,DocTrace始终优于现有开源基线和专有MLLM;与Qwen3-VL-8B-Instruct骨干模型相比,DocTrace在这三个基准上分别实现了14.4、11.3和11.7个点的绝对提升。除了具有竞争力的性能外,DocTrace还构建了具有显式节点级来源的可追溯证据图,为长文档理解提供了透明且可验证的推理。
英文摘要
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded evidence is progressively composed during reasoning, limiting both answer accuracy and traceability. In this paper, we cast LongDocVQA as an explicit evidence graph reasoning problem rather than implicit answer prediction. To this end, we propose DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance. To effectively learn these capabilities, we develop a two-stage training framework: joint Supervised Fine-Tuning (SFT) first initializes evidence localization and graph reasoning abilities, followed by task-specific Group Relative Policy Optimization (GRPO) with dedicated rewards to further optimize these capabilities. Extensive experiments on MMLongBench-Doc, LongDocURL, and SlideVQA demonstrate that DocTrace consistently outperforms both existing open-source baselines and proprietary MLLMs. Compared with the Qwen3-VL-8B-Instruct backbone, DocTrace achieves absolute improvements of 14.4, 11.3, and 11.7 points on the three benchmarks, respectively. Beyond competitive performance, DocTrace constructs traceable evidence graphs with explicit node-level provenance, enabling transparent and verifiable reasoning for long document understanding.