InSight-doc:面向长文档理解的智能体视觉感知框架
InSight-doc: Agentic Visual Perception for Long-Document Understanding
- The Hong Kong University of Science and Technology(香港科技大学)
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出智能体视觉感知框架InSight-doc,通过自适应选择放大高分辨率区域处理长文档,经SFT+RL训练后在文档VQA任务中显著提升准确率、降低幻觉与推理延迟,相关资源已公开。
AI中文摘要:
长文档理解通常需要对大量视觉丰富的页面进行推理,这使得推理成本高昂且易出现上下文退化问题。本研究提出InSight-doc,一种将视觉分辨率视为自适应推理时资源的智能体视觉感知框架,它从低分辨率开始,选择性地放大高分辨率区域以获取更精细的证据,无需依赖任何外部检索器。为训练该智能体,我们构建了包含1.79万个高质量SFT示例(带有区域级放大轨迹)和1.92万个困难RL示例的主动感知语料库。通过SFT+RL训练,InSight-doc-8B在文档VQA基准上较基线提升了4.3至16.4个准确率点;在长文档任务中,它在保持准确率优势的同时,将幻觉减少了40%以上,推理延迟降低了41%至68%。我们的代码、数据集和模型已在此httpsURL发布。
英文摘要:
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. InSight-doc starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct an active-perception corpus of 17.9K high-quality SFT examples with region-level zoom-in trajectories, accompanied by 19.2K hard RL examples. Through SFT+RL, InSight-doc-8B improves the baseline by 4.3--16.4 accuracy points over document VQA benchmarks. On long documents, it reduces hallucination by more than 40% and inference latency by 41%--68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m-Just/InSight-doc .