WholeBodyWAM:利用可扩展运动先验学习全身世界动作模型
WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors
浏览论文内容
中文总结 AI 辅助
WholeBodyWAM利用大规模异构运动数据预训练全身运动先验,通过非对称混合Transformer集成运动、视觉与动作专家,提升人形机器人全身操作的世界模型预测与数据效率。
中文摘要 AI 辅助
人形机器人全身操作需要协调的全身动力学,然而从目标机器人收集的大规模轨迹成本高昂且难以扩展。相比之下,来自人类和人形来源的全身运动数据非常丰富,尽管这些数据不能直接用作特定具身的机器人动作。本研究探讨这些可扩展的运动资源能否为人形世界动作建模提供可迁移的预测先验。我们提出了WholeBodyWAM,一种人形世界动作模型,它在目标机器人训练之前从大规模异构运动中学习全身动力学。我们构建了UniMotion-4K,一个涵盖超过4000小时来自人类视频、原生3D运动数据集和异构人形平台的运动语料库,并将这些多样来源规范化为统一的运动空间。随后,一个语言条件的运动专家被预训练以预测未来的全身运动,无需目标机器人的动作监督。在机器人后训练阶段,预训练的运动专家通过非对称混合Transformer(MoT)注意力与视觉专家和动作专家集成,使预测的场景动力学和全身运动能够共同为特定具身的动作生成提供信息。实验表明,WholeBodyWAM持续受益于增加的运动预训练规模,改善了未来运动预测和下游任务性能,并有效迁移到真实世界的人形操作中。此外,预训练的运动先验在有限的目标机器人演示下显著提高了数据效率。
英文摘要
Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.
发表机构
- Nankai University(南开大学)
- Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
- Beijing Institute of Technology(北京理工大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。