arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知晓并非总能表达:视觉语言模型中空间编码何时能抵达答案?

Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?

Zeyu Wang, Xinming Xu

arXiv 2608.22916首次发表:更新:

发表机构

Peking University; Tsinghua University(北京大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对视觉语言模型的空间编码何时能抵达答案的问题,采用方向干预方法,揭示了因果影响仅出现在中深层、文本思维链与视觉接地提示的不同作用等传输模式,为解决编码-接地差距提供了新视角。

AI 中文摘要

视觉语言模型(Vision-Language Models,VLMs)以隐藏状态编码空间信息,但回答问题时却常无法利用该信息。目前仍不清楚这些编码信息何时、何地能抵达答案。本研究通过方向干预(direction patching)解决该问题,这是一种应用于各层、 token 位置及提示格式的类别条件因果干预方法。利用基于先前编码证据构建的空间-ID方向,我们发现对答案 logit 的因果影响仅出现在中深层。文本思维链(chain-of-thought)会抑制多数模型中直接的物体词 argmax 级传输,而视觉接地提示则维持该传输。目标 logit 的正增益可保持在 argmax 阈值以下,且传输可在最终前缀 token 或更深层的答案步骤中重新出现。在研究的10个VLMs中,这些局部效应形成了可描述的传输模式,补充实验进一步表征了这些模式如何在数据集、属性及编码幅度间变化。综上,这些结果将编码-接地差距重新定义为VLMs中的条件传输问题。

英文摘要

Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this with direction patching, a class-conditioned causal intervention applied across layers, token positions, and prompt formats. Using spatial-ID directions constructed following prior encoding evidence, we find that causal influence on answer logits emerges only at mid-to-deep depths. Text chain-of-thought suppresses immediate object-word argmax-level transport in most models, while visually grounded prompts keep it open. Positive target-logit gain can remain below the argmax threshold, and transport can re-emerge at the final prefix token or at the answer step in deeper layers. Across the ten VLMs we study, these local effects form descriptive transport patterns. Complementary experiments characterize how these patterns shift across datasets, attributes, and encoding amplitudes. Together, these results reframe the encoding-grounding gap as a problem of conditional transport in VLMs.

CommentsAccepted to appear in the EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑