发表机构
ETH Zürich; MPI-IS; University of Tuebingen; ELLIS Institute; Vesoma(苏黎世联邦理工学院; 马克斯·普朗克智能系统研究所; 蒂宾根大学; 欧洲学习与智能系统研究所; 维索马公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VioLA是一种从人类数据学习的通用人形机器人控制策略,通过预测运动潜变量解决动作空间耦合与演示稀缺问题,零样本实现高成功率的运动与操作任务,且方法可适配多种模型主干。
AI 中文摘要
教人形机器人用全身遵循指令面临两个障碍:其动作空间庞大且耦合紧密,腿部、手臂和手指需协同运动同时保持平衡,导致关节级动作难以学习;且人形机器人演示数据稀缺,当前通用人形机器人策略无法直接遵循新指令,部署前需针对每个任务的远程操作演示进行微调。人类演示数据数量远多得多,但人的运动并非机器人指令。我们通过改变通用策略的预测内容来消除这两个障碍,提出VioLA——一种通用人形机器人策略,它预测身体和手部运动潜变量而非关节指令,预训练的身体与手部控制器会在机器人上执行这些潜变量,对应的运动编码器将人类和机器人的运动映射到相同的潜空间,因此人类记录会被标记为策略的动作空间,训练演示池包含1.406亿帧,其中93.2%来自人类。结果显示,VioLA在真实机器人上零样本遵循运动指令,无需任务特定微调,在GR00T N1.7和Ψ0分别达到16.7%和0%成功率的情况下,VioLA达到100%成功率;在无需任务特定微调的情况下,其操作成功率也达到88.6%。该方法在两种VLA和一种世界动作模型主干上均有效,仅基于人类演示训练的通用策略即可在真实机器人上零样本执行运动任务,代码和检查点将被发布。
英文摘要
Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.