看向何处?:视觉语言模型中视觉编码器的因果追踪
Where To Look? : Causal Tracing of Vision Encoders in VLM
浏览论文内容
中文总结 AI 辅助
本研究通过因果追踪方法,发现视觉语言模型的高因果性视觉标记常位于目标区域外,其强大多模态性能不代表存在空间定位因果表示,还揭示了视觉感知、利用与推理视觉结构的差距,提供了研究视觉信息处理的因果框架。
中文摘要 AI 辅助
视觉语言模型(Vision-Language Models)能以极高的准确率描述图像,但一个更根本的问题仍未得到解答:究竟是哪些视觉信息驱动了它们的回答?本研究通过因果追踪(causal tracing)对此展开调查,发现高因果性的视觉标记(vision tokens)往往位于目标区域之外。将分析扩展到更大规模的视觉语言模型后,在不同模型和干扰设置下均观察到类似模式,这表明强大的多模态性能并不一定意味着存在空间定位的因果表示。我们进一步探究:当外观线索被移除时,这些模型能否保留视觉结构?结果发现,模型会利用视觉线索来理解视觉结构。综合实验结果,本研究揭示了视觉感知、利用及推理视觉结构之间存在的差距,并提供了一个因果框架,用于研究现代视觉语言模型如何对视觉信息进行转换、保留和最终利用。
英文摘要
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
发表机构
- Indian Institute of Technology Gandhinagar(印度甘地纳格尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。