arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DocAtlas:将长文档理解视为可变状态交互

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining Guo

arXiv 2608.07527首次发表:更新:

发表机构

Wuhan University; Microsoft(武汉大学; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出DocAtlas系统,将长文档理解视为可变状态交互,通过可变文档工具结合自改进检索等技术,在MMLongBench-Doc上用GPT-5.4达71.4%,还可提升紧凑型VLM智能体性能。

AI 中文摘要

长文档理解需要模型跨多个页面、布局、表格、图形和图表查找并整合证据。现有的检索增强系统通常在生成前从静态索引中选择证据,而近期的智能体系统虽增加了多轮工具使用,但往往依赖行为由提示设定的冻结专有主干模型。我们提出DocAtlas,一种将长文档理解视为可变状态信息寻求过程的系统。我们将DocAtlas实例化为可变文档工具:一个外部环境,用于决定每一步搜索、读取、存储、审查和向模型展示的文档信息。给定文档和问题,该工具提供搜索、读取、记笔记和审查工具,维护分层树和笔记存储,并在智能体记录证据时更新两者。DocAtlas在固定上下文预算下结合自改进检索、选择性证据访问和主动工作记忆。同一工具支持大型VLMs的推理时使用,以及针对紧凑型VLM智能体的端到端强化学习。使用GPT-5.4,DocAtlas在MMLongBench-Doc上达到71.4%,超过人类专家参考的65.8%。在DocAtlas环境中经端到端RL训练的Qwen3.5-4B VLM达到63.7%,相比直接输入基线的54.4%,表明可变文档工具设计可大幅提升紧凑型文档智能体的性能。

英文摘要

Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen proprietary backbones whose behavior is set by prompts. We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process. We instantiate DocAtlas as a mutable document harness: an external environment that determines what document information is searched, read, stored, reviewed, and shown to the model at each step. Given a document and question, the harness exposes search, reading, note-taking, and review tools, maintains a hierarchical tree and note store, and updates both as the agent records evidence. DocAtlas combines self-improving retrieval, selective evidence access, and active working memory under a fixed context budget. The same harness supports inference-time use with large VLMs and end-to-end reinforcement learning for compact VLM agents. With GPT-5.4, DocAtlas reaches 71.4\% on MMLongBench-Doc, exceeding the human-expert reference of 65.8\%. A Qwen3.5-4B VLM trained with end-to-end RL in the DocAtlas environment reaches 63.7\%, compared with a 54.4\% direct-input baseline, showing that mutable document-harness design can improve compact document agents by a large margin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑