检索头与视觉结合:揭示视觉语言模型如何定位和提取视觉信息
Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information
浏览论文内容
中文总结 AI 辅助
该研究发现视觉语言模型(VLMs)中存在视觉检索头(VRHs),其占注意力头的1.7%-2.6%,负责将文本描述与图像区域建立因果关联,屏蔽前20个VRHs可使定位准确率降低多达80个百分点,且VRHs具备泛化性、功能特异性与架构共享性。
中文摘要 AI 辅助
视觉语言模型(VLMs)能够定位文本提示所指代的图像区域,并将对应的视觉证据传递至输出端,但这一行为背后的内部机制尚不明确。受大型语言模型中检索头的启发,我们探究VLMs是否包含类似的视觉检索机制。我们通过引入视觉检索头(VRHs)给出肯定答案,VRHs是注意力头的一小部分(约1.7%-2.6%),负责将文本描述与图像区域建立因果关联。为找到VRHs,我们在查询 token、键聚合和跨样本聚合的统一设计空间下重构了现有的头评分方法。随后我们证明,通过对真实指代区域求和来评分输出预测 token 的注意力,能最可靠地识别因果头。在11个VLMs和5个指代表达基准上,仅屏蔽排名前20的VRHs会使定位准确率降低多达80个百分点,而屏蔽相同数量的随机头几乎没有影响。除了复现文本检索头确立的因果性、稀疏性、通用性三重特性外,VRHs还表现出此前未被报道的多种属性:它们可在视觉指代任务间泛化,尽管是通过边界框预测发现的,但在属性、空间、计数和视觉数学基准上仍保持因果性;它们具有功能特异性,在破坏定位的同时保留输出格式;且具有架构共享性,在共享大语言模型(LLM)主干但视觉编码器、投影器和指令微调不同的VLMs间可传递因果性。
英文摘要
Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under a unified design space over query tokens, key aggregation, and cross-sample aggregation. We then show that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads. Across eleven VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect. Beyond replicating the causal-sparse-universal triad established for text retrieval heads, VRHs exhibit several properties not previously reported: they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction; they are functionally specific, preserving output format while corrupting localization; and they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。