arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

grounding 不等于知晓:视觉语言模型是否需要对象定位来进行空间推理?

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li, Zhixiang Lu, Haolin Yang, Imran Razzak, Yutong Xie

arXiv 2608.23074首次发表:更新:

AI 中文总结

本研究以 LLaVA-1.5 和 Qwen2.5-VL 为对象,通过多种可解释性工具揭示 VLMs 空间推理的机制,发现其空间关系预测无需精确对象定位,定位与空间推理共享早期通路但依赖部分不同通路。

AI 中文摘要

视觉语言模型(VLMs)能够回答空间问题,但将对象定位与空间推理关联的机制仍未被充分理解。目前尚未深入探究空间推理是否在内部需要精确的对象定位,还是可以通过全局布局线索绕过显式定位。本研究使用一系列机制可解释性工具,包括 token 消融、分层探测、注意力敲除和因果中介分析,对两个代表性模型家族 LLaVA-1.5 和 Qwen2.5-VL 展开研究。研究发现,空间关系预测遵循分阶段的定位到推理过程,其中与对象对齐的 token 建立了粗略的目标-参考锚点,而不需要精确的边界框边界。位置信息在关系决策出现前即可被解码,且有一小部分注意力头介导了定位和空间推理的因果效应。这两项任务共享早期与定位相关的处理过程,但最终依赖于部分不同的专用通路。通过严谨的实验,本研究在 token、层和头层面阐明了 VLMs 如何将对象定位转化为空间关系,表明知晓对象的位置并不等同于知晓它们之间的关系。

英文摘要

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires precise objects localization, or can bypass explicit localization through global layout cues. In this work, we investigate two representative model families, LLaVA-1.5 and Qwen2.5-VL, using a suite of mechanistic interpretability tools, including token ablation, layer-wise probing, attention knockout, and causal mediation analysis. We find that spatial relation prediction follows a staged grounding-to-reasoning process in which object-aligned tokens establish coarse target-reference anchors, while precise bounding-box boundaries are not required. Positional information becomes decodable before relation decisions emerge, and a small set of attention heads mediates the causal effects of both localization and spatial reasoning. The two tasks share early grounding-related processing but ultimately rely on partially distinct specialized pathways. Through rigorous experiments, we provide a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑