发表机构
New York University; Reflex; Université de Montréal; Maastricht University(纽约大学; Reflex; 蒙特利尔大学; 马斯特里赫特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究动作后训练对VLM空间理解的影响,发现VLA存在深层解码能力的“基底”差距与“悬崖”骤降,定位其源于深层MLP干扰,消融该模块可恢复多数解码性能。
AI 中文摘要
在构建视觉-语言-动作模型(VLA)的动作后训练过程后,视觉语言模型(VLM)保留了多少空间理解能力?我们在权重匹配的开源基础VLM/VLA对(Molmo2-ER和MolmoAct2-LIBERO)的每个解码器层上探测空间几何理解的基础——深度感知。首先,VLA在所有层的深度解码效果更差,这一持续差距我们称为“基底”。其次,退化并非均匀:基础VLM的深度解码能力在最后几层有所提升,而VLA的则出现骤降,我们将这一额外的深层骤降称为“悬崖”。我们通过因果定位发现,该悬崖源于深层多层感知机(MLP)的干扰:消融深层MLP权重可恢复大部分终端解码悬崖,而匹配的注意力消融及权重匹配基础VLM中的相同干预均未产生可比恢复。模块级分解解释了这种差异:基础VLM的深度最易通过累积的MLP权重获取,而动作后训练会在后期累积权重中破坏深度解码能力。
英文摘要
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM's depth decodability improves through its final layers, the VLA's collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.
CommentsAccepted to the archival proceedings track of the Embodied Multimodal Reasoning (EMR) Workshop at ECCV 2026