发表机构
University of Edinburgh; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(爱丁堡大学; 中国科学院深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SAVLA,一种对称性感知的VLA模型,通过冻结视觉-语言骨干、结合等变流匹配动作头和规范化器,在LIBERO上平均成功率提升5.1点,旋转下成功率从41.5%提升至90.4%。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为语言条件机器人操作的主导范式。然而,尽管图像和语言指令本质上编码了几何信息,VLA模型的空间能力完全是从演示中获得的。因此,它们仅在演示所覆盖的场景姿态范围内可靠。我们提出了SAVLA,一种端到端的对称性感知VLA模型,用于鲁棒且数据高效的政策学习。我们的方法保持预训练的视觉-语言骨干完全冻结,同时将其与等变流匹配动作头和学习的规范化器相结合。该头将其状态、动作和条件输入分解为不变和等变通道,并在其所有层中保持这种类型化。规范化器将斜视图像转换为规范框架,并一致地旋转几何条件。我们在LIBERO上评估了我们的模型。与GR00T N1.5基线相比,SAVLA在四个LIBERO套件上的平均成功率提高了5.1个百分点,并将LIBERO-Goal在旋转下的平均成功率从41.5%提高到90.4%。
英文摘要
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the range of scene poses that the demonstrations cover. We propose SAVLA, an end-to-end symmetry-aware VLA model for robust and data-efficient policy learning. Our approach keeps the pretrained vision-language backbone entirely frozen while combining it with an equivariant flow-matching action head and a learned canonicalizer. The head decomposes its state, action, and conditioning inputs into invariant and equivariant channels, and preserves this typing throughout all of its layers. The canonicalizer transforms oblique-view images into a canonical frame and rotates the geometric conditions consistently. We evaluate our model on LIBERO. Compared with the GR00T N1.5 baseline, SAVLA improves the success rate averaged over all four LIBERO suites by 5.1 points and increases the mean success rate under rotation on LIBERO-Goal from 41.5% to 90.4%.
Comments8 pages, 4 figures, 6 tables