arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从恢复到骤降:动作后训练如何降低视觉语言模型的深层解码能力

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou

arXiv 2608.08904首次发表:更新:

发表机构

New York University; Reflex; Université de Montréal; Maastricht University(纽约大学; Reflex; 蒙特利尔大学; 马斯特里赫特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究动作后训练对VLM空间理解的影响,发现VLA存在深层解码能力的“基底”差距与“悬崖”骤降,定位其源于深层MLP干扰,消融该模块可恢复多数解码性能。

AI 中文摘要

在构建视觉-语言-动作模型(VLA)的动作后训练过程后,视觉语言模型(VLM)保留了多少空间理解能力?我们在权重匹配的开源基础VLM/VLA对(Molmo2-ER和MolmoAct2-LIBERO)的每个解码器层上探测空间几何理解的基础——深度感知。首先,VLA在所有层的深度解码效果更差,这一持续差距我们称为“基底”。其次,退化并非均匀:基础VLM的深度解码能力在最后几层有所提升,而VLA的则出现骤降,我们将这一额外的深层骤降称为“悬崖”。我们通过因果定位发现,该悬崖源于深层多层感知机(MLP)的干扰:消融深层MLP权重可恢复大部分终端解码悬崖,而匹配的注意力消融及权重匹配基础VLM中的相同干预均未产生可比恢复。模块级分解解释了这种差异:基础VLM的深度最易通过累积的MLP权重获取,而动作后训练会在后期累积权重中破坏深度解码能力。

英文摘要

How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM's depth decodability improves through its final layers, the VLA's collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.

CommentsAccepted to the archival proceedings track of the Embodied Multimodal Reasoning (EMR) Workshop at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑