arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉-语言-动作模型是否理解并适应它们操作的物体,还是仅仅重放已学行为?

Do VLAs Understand and Adapt to the Objects They Handle, or Simply Replay Learned Behaviors?

Xinnuo Xu

arXiv 2610.06078首次发表:更新:

AI 中文总结

本文通过线性探测和案例研究,发现VLA模型对物体物理属性编码较弱,且不能可靠地利用这些属性调整动作,其泛化更多源于偶然鲁棒性而非真正理解。

AI 中文摘要

本文探讨视觉-语言-动作(VLA)模型的泛化能力是否根植于对物体物理属性的全局理解,从而使其能够针对未见过的设置调整动作,还是模型仅仅重放那些在新设置中恰好成功的已学动作。前者反映真正的泛化,后者反映偶然的鲁棒性。我们首先通过线性探测和表征相似性分析(RSA)对七个VLA模型的激活进行检验,考察其对物理属性的感知。我们发现,在几乎所有模态流中,质量、易碎性、可变形性、摩擦力和尺寸等物理属性的可解码性低于语义类别、材料、声音和价格等非物理属性。与基础视觉-语言模型(VLM)相比,机器人预训练削弱了语言流中物理属性的线性编码。无论是预训练还是下游微调,都没有加强物理属性差异与激活距离之间的对齐。随后,我们探究这些激活中存在的微弱物理信息是否影响VLA生成的动作。在一个受控的LIBERO案例研究中,我们增加了域内物体的质量,并通过语言或视觉传达这一变化。大多数VLA对较重物体和原始质量物体采用相似的提升行为,导致任务成功率下降。少数例外模型会响应词汇或视觉线索而非质量本身来改变行为。这些结果表明,VLA对物理属性的编码较弱,且不能可靠地利用它们来调整动作。

英文摘要

This paper asks whether VLA generalization is grounded in a global understanding of objects' physical properties that enables policies to adapt their motion to unseen setups, or if policies simply replay the motions they've learnt that happen to succeed in new setups. The former reflects genuine generalization; the latter reflects incidental robustness. We first examine awareness of physical properties in seven VLAs by applying linear probing and representational similarity analysis (RSA) to their activations. We find that physical properties, including mass, fragility, deformability, friction and size are less decodable than non-physical properties such as semantic category, material, sound and price in nearly every modality stream. Compared with their base VLMs, robot pre-training weakens the linear encoding of physical properties in the language stream. Neither pre-training nor downstream fine-tuning strengthens the alignment between physical-property differences and activation distances. We then ask whether the weak physical information present in these activations shapes the actions a VLA generates. In a controlled LIBERO case study, we increase the mass of an in-domain object and signal the change through language or vision. Most VLAs use similar lifting behaviour for the heavier and original-mass objects, leading to task success declines. The few exceptions change their behaviour in response to lexical or visual cues rather than to mass itself. These results suggest that VLAs encode physical properties weakly and do not reliably use them to adapt their motion.

Comments8 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑