EviViT: 证据自适应视觉 Transformer 用于细粒度感知
EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception
- Tsinghua University(清华大学)
- Peng Cheng Laboratory(鹏城实验室)
- The Hong Kong University of Science and Technology(香港科技大学)
- LMMs-Lab
- Beijing Institute of Technology(北京理工大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EviViT 是一种轻量级附件,通过人类视觉搜索轨迹监督的证据密度引导视觉 Transformer 在细粒度感知中按需获取细节,并在冻结骨干下提升多宿主准确率且节省视觉标记。
AI中文摘要:
细粒度视觉感知使视觉语言模型能够区分细微属性并将其答案基于视觉证据。在高分辨率场景中,以更高分辨率处理整个图像会将视觉标记花费在无关内容上,而孤立的裁剪可能会丢失解释所选证据所需的上下文。我们引入了 EviViT,这是一种轻量级附件,它学习预训练视觉 Transformer 应在何处获取细节。人类视觉搜索轨迹监督一个基于问题的证据密度,该密度指导从原始像素进行区域重读以及视觉标记的分配。然后,一个稀疏的、坐标感知的桥接将区域特征连接到全局场景,使宿主能够在上下文中解释精确的证据。在宿主骨干冻结的情况下学习,该附件无需重新拟合即可服务于基础模型和兼容的后训练后代。在九个宿主上的实验显示平均细粒度准确率持续提升。匹配预算的比较进一步表明,EviViT 在每个测试的标记上限下都优于仅全局处理,同时使用更少的视觉标记。
英文摘要:
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.