State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models
面向文档视觉语言模型中细粒度感知的状态条件视觉证据检索
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; ByteDance, LarkAI(字节跳动 LarkAI) ; Shanghai Key Lab of Intelligent Information Processing(上海智能信息处理重点实验室)
专题命中 文档图表理解 :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.CV
AI总结 针对现有VLM文档解析方法效率低的问题,提出SCVER方法,通过状态条件检索高分辨率区域,结合SGLO稳定训练,提升了低输入分辨率下的鲁棒性与准确率-效率权衡。
Comments 18 pages, 12 figures