arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05023cs.CV

看向你所说之处:用于视觉推理的自接地注意力

Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning

Uri Berger, Gal Chechik, Gal Dalal

首次发表
浏览论文内容

中文总结 AI 辅助

提出自显著性方法,利用接地模型定位推理提及区域来监督VLM视觉注意力,在25个基准上显著优于先前方法,并揭示注意力边界偏差。核心贡献是条件化目标区域于生成推理以提升性能。

中文摘要 AI 辅助

我们引入了自显著性(Self-Saliency),这是一种训练视觉语言模型(VLMs)的方法,旨在提高其视觉注意力与其推理中提到的图像区域之间的一致性。自显著性使用一个接地模型来定位每个推理步骤中提到的对象,并将所得区域作为模型视觉注意力的监督。以往关于引导视觉注意力的工作仅基于图像和问题来确定目标图像区域。相比之下,我们表明,将目标区域基于模型生成的推理进行条件化可以提高下游性能。为了进行适当评估,我们构建了一个统一的、广泛的25个视觉推理基准套件,在其中复现了以往方法的结果。我们发现,自显著性在平均排名和平均分数上都显著优于先前的注意力引导方法和基于图像级文本接地的基线方法。训练后分析表明,模型主要调整其推理文本以适应现有的注意力模式,产生更短的步骤并引用更大的区域。然而,在控制生成的文本时,对已接地区域的注意力在相关层中显著增加。最后,我们识别出VLM视觉注意力中朝向图像边缘的一致几何偏差。然而,我们的消融实验表明,自显著性的增益不能简单地通过将注意力与图像中心对齐来解释,这突显了将视觉注意力与模型推理中提到的区域对齐的重要性。

英文摘要

We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.

发表机构

  • NVIDIA Research(英伟达研究院)
  • The Hebrew University of Jerusalem(耶路撒冷希伯来大学)
  • University of Melbourne(墨尔本大学)
  • Bar-Ilan University(巴伊兰大学)

机构由 AI 辅助整理,请以论文原文为准。

↑