arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28698cs.CV

面向文档视觉语言模型中细粒度感知的状态条件视觉证据检索

State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

Mingxu Chai, Chenyu Liu, Ziyu Shen, Jiazheng Zhang, Kaidi Zhang, Ruoyu Chen, Jun Long, Jihua Kang, Tao Gui, Qi Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有VLM文档解析方法效率低的问题,提出SCVER方法,通过状态条件检索高分辨率区域,结合SGLO稳定训练,提升了低输入分辨率下的鲁棒性与准确率-效率权衡。

中文摘要 AI 辅助

与典型的视觉语言任务相比,文档解析对细粒度视觉感知提出了更高要求。现有基于视觉语言模型(VLM)的解析方法依赖全局压缩的视觉标记,其中细粒度细节被纠缠在单一表示中,并在解码过程中被反复访问。然而,我们发现,每次预测所需的视觉证据通常是局部的,且受当前解码状态的约束,而这类表示必须在每个解码步骤中完整访问,导致计算效率低下。为解决这种不匹配,我们将感知建模为自回归解码过程中的状态条件视觉证据检索(SCVER)。该模型基于紧凑的全局表示处理粗结构,并根据当前标记状态检索一小部分相关的高分辨率区域。这种由粗到细的设计支持按需访问细粒度视觉线索,减轻全局共享表示对所有细粒度细节进行编码的负担。我们进一步发现,在视觉语言模型中学习此类状态条件检索具有挑战性且不稳定。为稳定该过程,我们引入空间引导学习目标(SGLO)来指导检索过程。在文档解析基准上的实验表明,SCVER在输入分辨率降低的情况下提高了鲁棒性,并实现了更优的准确率-效率权衡,证明了按需视觉证据检索对细粒度感知的有效性。

英文摘要

Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approaches rely on globally compressed visual tokens, where fine-grained details are entangled within a single representation and repeatedly accessed during decoding. However, we observe that the visual evidence for each prediction is typically localized and conditioned on the current decoding state, whereas such representations must be accessed in full at every decoding step, resulting in inefficient computation. To address this mismatch, we formulate perception as state-conditioned visual evidence retrieval (SCVER) during autoregressive decoding. The model operates on a compact global representation for coarse structure and retrieves a small set of relevant high-resolution regions conditioned on the current token state. This coarse-to-fine design enables on-demand access to fine-grained visual cues, relieving globally shared representations from encoding all fine-grained details. We further find that learning such state-conditioned retrieval in VLMs is challenging and unstable. To stabilize this process, we introduce a Spatially-Guided Learning Objective (SGLO) to guide the retrieval process. Experiments on document parsing benchmarks show that SCVER improves robustness under reduced input resolution and achieves a better accuracy-efficiency trade-off, demonstrating the effectiveness of on-demand visual evidence retrieval for fine-grained perception.

发表机构

  • College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
  • Shanghai Innovation Institute(上海创新研究院)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • ByteDance, LarkAI(字节跳动 LarkAI)
  • Shanghai Key Lab of Intelligent Information Processing(上海智能信息处理重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑