发表机构
University of Illinois Urbana-Champaign; Texas A&M University; University of California, Irvine(伊利诺伊大学厄巴纳-香槟分校; 德州农工大学; 加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对VLA策略微调中存在的问题提出Anchor-Align方法,通过视觉语言锚定和语言-动作对齐两个目标增强BC,在物理机器人和模拟测试中均有效果提升,表明保留预训练表示与有效动作学习可兼顾。
AI 中文摘要
通过行为克隆(BC)在机器人演示上对预训练的视觉语言模型(VLM)进行微调已成为视觉语言动作(VLA)策略的标准方法。然而,BC微调会逐步覆盖支持视觉和语义泛化的预训练表示。在网络图像-文本数据上进行联合训练并不能防止这种情况,它将语言和动作损失应用于单独的观察,导致VLA存在语言-动作不对齐,而标准操作基准并未暴露这一点。我们提出了Anchor-Align方法,它通过两个目标增强BC:视觉语言锚定从冻结的VLM副本中提取分层表示以防止这种漂移,而语言-动作对齐将每个动作目标转换为离散的运动方向标签,并在相同的机器人观察上联合训练语言和动作预测。在物理xArm7机器人上,在两种广泛使用的VLA架构中,Anchor-Align都提高了实际机器人的成功率(分别从28%提高到54%和从37%提高到60%)。在模拟中大规模测试时,我们分别在LIBERO-PRO、LIBERO-Plus和CALVIN上展示了在分布外扰动、感知鲁棒性和长视野控制方面的持续改进,这表明保留预训练表示和有效的动作学习在根本上并不矛盾。
英文摘要
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, applies language and action losses to separate observations, leaving VLAs with language-action misalignment that standard manipulation benchmarks do not expose. We propose Anchor-Align, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation. We conduct real-world evaluations across eight manipulation settings on single-arm xArm7 and bimanual YAM robots, using two VLA architectures with regression and flow-matching action heads. Across these settings, Anchor-Align consistently improves over BC on novel targets, layouts, and motion-sensitive bimanual tasks requiring coordinated control. At scale in simulation, we demonstrate consistent improvements on OOD perturbations, perceptual robustness, and long-horizon control across LIBERO-PRO, LIBERO-Plus, and CALVIN, respectively, suggesting that preserving pretrained representations and effective action learning are not fundamentally at odds. Project page: anchoralignvla.github.io
CommentsCode: https://github.com/dwipddalal/Anchor-Align