arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉-语言-动作模型理解指令吗?关于语言接地机制的机械可解释性研究

Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding

Theodor Wulff, Angelo Cangelosi

arXiv 2610.10178首次发表:更新:

发表机构

The University of Manchester(曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过激活和归因修补,系统分析π0.5和GR00T N1.7对语言指令的依赖,发现两模型对方向性语言敏感,但敏感性位置和归因可靠性因模型而异。

AI 中文摘要

视觉-语言-动作模型旨在跨环境和任务描述进行泛化,这引发了一个问题:它们的动作生成是否真正依赖于语言指令,还是主要依赖视觉线索和表面相关性。对视觉和语言观测空间中的方差具有鲁棒性对于实际部署至关重要,然而VLA模型缺乏显式的接地模块,而是依赖其视觉-语言模型骨干的内在语言接地能力。为此,我们对两个最先进的视觉-语言-动作模型π0.5和GR00T N1.7的语言接地能力进行了受控的机械可解释性研究,通过对动作生成模块的残差流应用激活和归因修补。我们按照五种策略系统地破坏LIBERO基准输入样本的任务指令:同义词替换、语义缩放、方向性破坏、随机物体替换和空字符串。我们的实验发现,两个模型对抽象改写和引用不存在的物体相对不敏感,但对空任务描述反应强烈,尤其是对方向性语言。在动作生成过程中,这种敏感性集中在每个模型的不同位置:对于GR00T N1.7,主要集中在前期的周期性交叉注意力层,而对于π0.5,则分布在最早层和选定的后期层。对于GR00T N1.7,方向性扰动驱动了一些最大的因果效应,同时内部表征几何结构相对不变,这种分离在π0.5中并未清晰观察到。最后,归因修补的可靠性依赖于模型:对于GR00T N1.7,它与激活修补紧密吻合,但对于π0.5则不然。

英文摘要

Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations. Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones. For this reason, we conduct a controlled mechanistic interpretability study on the language grounding capabilities of two state-of-the-art Vision-Language-Action models, $π_{0.5}$ and GR00T N1.7, by applying activation and attribution patching to the residual stream of the action generation modules. We systematically corrupt the task instruction of input samples of the LIBERO benchmark following five strategies: synonym replacement, semantic scaling, directional corruption, random object substitution, and empty string. Our experiments find that both models are comparatively insensitive to abstract rephrasing and to referencing non-existent objects, but react strongly to empty task descriptions and, especially, to directional language. During action generation, this sensitivity is concentrated in different loci for each model: mainly in the early, periodic cross-attention layers for GR00T N1.7, versus distributed across the earliest and selected later layers for $π_{0.5}$. For GR00T N1.7, directional perturbations drive some of the largest causal effects while leaving the internal representational geometry comparatively unchanged, a dissociation we do not observe clearly for $π_{0.5}$. Finally, the reliability of attribution patching is model-dependent: it closely tracks activation patching for GR00T N1.7 but not for $π_{0.5}$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑