arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30092cs.ROcs.CV

自适应视觉-语言-动作模型用于鲁棒机器人部署

Self-Adaptive VLA for Robust Robot Deployment

Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang, Zhenjia Xu, Chuang Gan

首次发表
浏览论文内容

中文总结 AI 辅助

提出自适应VLA训练后配方,利用自身轨迹作为上下文,通过上下文编码器和自适应层归一化,使策略在硬件变化下自我适应,恢复80%以上性能,提升部署鲁棒性。

中文摘要 AI 辅助

尽管视觉-语言-动作(VLA)模型在机器人操作中展现出令人印象深刻的能力,但其无记忆特性使其在测试时的环境变化面前显得脆弱,尤其是由磨损或校准不完善引起的硬件变化。在部署过程中使这些模型能够自我适应,而无需持续进行现场重新校准,仍然是实现现实世界可扩展性的关键瓶颈。在这项工作中,我们引入了自适应VLA,一种新颖的训练后配方,使策略能够利用自身的轨迹作为上下文,迭代地适应部署时的硬件变化。为此,我们首先在故意注入的硬件变化下收集策略轨迹。然后,我们通过针对这些已知变化对专家动作进行预补偿,将基础策略的训练数据转换为条件于变化的专家演示。接下来,我们引入一个轻量级的即插即用上下文编码器,将上下文(包括视觉观察、本体感觉和变化环境中的动作)压缩为潜在上下文令牌。该令牌通过自适应层归一化(AdaLN)调制策略。此外,我们发现上下文令牌可以集成,使策略能够逐步迭代地自我纠正并缓解失败。在四个精度关键的双臂和灵巧操作任务上的大量实验表明,自适应VLA在硬件变化(如驱动偏差和关节编码器偏移)下恢复了基础策略性能的80%以上。此外,与基础策略相比,自适应VLA使得在新工作站上的部署更加鲁棒。我们的方法为鲁棒的大规模现实世界机器人部署和更简单的维护提供了一条途径。视频见此https URL。

英文摘要

While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy's training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy's performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.

发表机构

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Genesis AI

机构由 AI 辅助整理,请以论文原文为准。

↑