HuMBLE:基于人体运动驱动的具身运动行为学习
HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion
浏览论文内容
中文总结 AI 辅助
本文提出HuMBLE框架,通过教师-学生蒸馏和多任务RL平衡指令跟踪与人体数据风格,生成可操控、鲁棒且仿生的人形运动策略,并在多款机器人上验证。
中文摘要 AI 辅助
尽管近期人形机器人运动领域取得了进展,针对指令跟踪和鲁棒性进行优化的控制器往往会产生机械化的步态,而与人体运动数据绑定的控制器则常常无法泛化到数据分布之外的指令。本研究引入了一个学习框架,以平衡这些相互竞争的目标,从人体数据中合成实时可操控、鲁棒且仿生的运动策略。利用一个涵盖多种速度和方向的内部整理的运动数据集,我们首先通过教师-学生蒸馏过程学习一个自然运动先验策略。具体而言,我们使用强化学习(RL)训练一个全身参考条件策略,然后将其蒸馏为一个仅以本体感觉和平面躯干速度转向指令为条件的轻量级先验策略。接下来,我们使用多任务强化学习对先验策略进行微调,以扩展指令覆盖范围和鲁棒性,超越数据分布,将跟踪任意指令的目标条件任务与跟踪人体数据作为显式风格正则化器的参考引导任务配对。我们在三款人形机器人上验证了我们的框架:波士顿动力公司的Atlas R1、Atlas D1和宇树科技的G1。实验结果表明,在现实世界场景中具有鲁棒性能,包括室内和室外环境中的直接用户控制运动,以及作为分层控制栈中运动层的集成。与未使用人体数据训练的Tabula Rasa RL策略的基准比较和消融研究证实,我们的框架产生了一个轻量级、可部署的策略,该策略能从转向指令重建协调的全身行为,在保持鲁棒性和完全可操控性的同时保留人体步态特征。
英文摘要
Despite recent advances in humanoid locomotion, controllers optimized for command tracking and robustness tend to produce mechanical gaits, whereas controllers tied to human motion data often fail to generalize to commands outside the data distribution. This work introduces a learning framework that balances these competing objectives to synthesize real-time steerable, robust, and biomimetic locomotion policies from human data. Using an in-house curated locomotion dataset covering diverse speeds and directions, we first learn a natural locomotion prior policy through a teacher-student distillation process. Specifically, we train a full-body reference-conditioned policy with Reinforcement Learning (RL), then distill it into a lightweight prior policy conditioned solely on proprioception and a planar torso-velocity steering command. Next, we fine-tune the prior policy with multi-task RL to expand command coverage and robustness beyond the data distribution, pairing a goal-conditioned task that tracks arbitrary commands with a reference-guided task that tracks the human data as an explicit style regularizer. We validate our framework on three humanoid robots: the Boston Dynamics Atlas R1, Atlas D1, and Unitree G1. Experimental results demonstrate robust performance across real-world scenarios, including direct user-controlled locomotion in indoor and outdoor environments, and integration as the locomotion layer within hierarchical control stacks. Benchmarks against Tabula Rasa RL policies trained without human data and ablation studies confirm that our framework yields a lightweight, deployable policy that reconstructs coordinated whole-body behavior from a steering command, retaining the human gait characteristics while remaining robust and fully steerable.
发表机构
- RAI Institute, USA(RAI研究所,美国)
- Boston Dynamics, USA(波士顿动力公司,美国)
- Carnegie Mellon University, USA(卡内基梅隆大学,美国)
机构由 AI 辅助整理,请以论文原文为准。