发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对VLA模型部署到新机器人时因硬件不匹配导致性能下降的问题,提出自监督方法生成在线交互rollout数据微调VLA模型,使模型在目标机器人上同时具备先验任务能力、指令遵循能力与高效学习新技能的能力。
AI 中文摘要
现有最先进的视觉-语言动作(VLA)模型如π₀.₅具备强大的语义理解、指令遵循及任务行为能力,但部署到新机器人时,若硬件配置与预训练存在微小不匹配,就会导致性能严重下降。在新实体机器人的领域内专家数据上微调VLA模型,虽能提升专家任务的性能,但会损失其原有的指令遵循能力与行为先验。本文提出一种自监督方法,从零样本VLA模型生成在线交互rollout数据作为额外训练数据用于微调。实验表明,该微调方案能生成强大的多任务策略,在目标机器人上:1)继承从零样本模型提炼的先验任务;2)具备通用指令遵循能力;3)从专家数据学习新技能且样本效率提升。我们在真实ALOHA机器人及RoboTwin新模拟基准的泛化测试集上验证了该方法的有效性,视频结果可在该网址获取。
英文摘要
State-of-the-art vision-language-action (VLA) models such as $π_{0.5}$ exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. Finetuning the VLA on in-domain expert data from the new embodiment improves performance on the expert task but leads to a loss in its original instruction following and behavioral priors. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning. Our experiments show this finetuning scheme yields strong multi-task policies that, on the target robot, (1) inherit prior tasks distilled from the zero-shot model, (2) enable generalist instruction following, while (3) learning new skills from expert data with improved sample efficiency. We demonstrate the success of our approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin. Video results are available at https://self-supervised-control.pages.dev/
CommentsProject Page: https://self-supervised-control.pages.dev/