arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18823cs.CVcs.AIcs.CL

使用OCR头来语言化图像语义

Using OCR Heads to Verbalize Image Semantics

  • Northeastern University(东北大学)
  • Independent(独立)

机构由 AI 辅助整理,请以论文原文为准。

Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau

AI总结:

本研究通过识别视觉语言模型中用于OCR的注意力头,发现其输出可解释语义特征,并利用这些头构建语言化透镜变换,揭示图像表示与语言在早期层对齐,且可通过逆变换编辑图像概念,为可解释性研究提供新视角。

AI中文摘要:

视觉语言模型(VLMs)如何从像素映射到语义?为了理解这个普遍问题,我们聚焦于一个狭窄的问题:研究VLMs如何执行光学字符识别(OCR)。在四个模型中,我们识别出对OCR因果必要的注意力头,并发现这些实际上是通用头,它们在所有图像标记上输出可解释的语义特征。例如,将这些头指向包含单词“bike”的图像标记会导致Qwen3-VL-8B输出“bike”,但将它们指向鸟翅膀则导致模型输出标记“feathers”。我们将这些头的注意力权重折叠成一个单一的语言化透镜变换,该变换揭示了所有层中隐藏状态的可解释语义特征。当与词汇空间投影结合时,我们可以从第0层开始获得可解释的标签,表明图像表示在早期层中实际上与语言对齐。我们发现我们还可以使用此变换的逆变换来编辑非单词概念,例如,在自然图像中用左轮手枪替换拖拉机,提供了因果证据表明该子空间不仅对OCR有用。我们的结果展示了特定机制的研究如何能够阐明更广泛的可解释性问题。

英文摘要:

How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word "bike" causes Qwen3-VL-8B to output "bike," but pointing them at a bird wing causes the model to output the token "feathers." We collapse these heads' attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.

补充信息

↑