从运动模仿中学习可重复使用的类人运动混合先验
Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation
浏览论文内容
中文总结 AI 辅助
研究提出三阶段流程,将运动模仿技能转化为可重复使用的类人运动混合先验(HMP)。先训练专家策略模仿人类运动,再提炼为冻结架构,最后训练任务级策略。经模拟评估和机器人验证,HMP可重复使用,还揭示可解释码本结构,训练码本技巧能减少下游跌倒。
中文摘要 AI 辅助
强化学习可生成强大的类人控制器,但每个新任务通常作为单独策略训练,有其自身奖励设计和训练过程。运动模仿通过训练策略跟踪重新目标化的人类运动提供了运动能力的替代来源,但由此产生的控制器仍是参考跟踪器,不能直接用作任务策略。我们提出一个三阶段流程,将运动模仿技能转化为可重复使用的类人运动混合先验(HMP)。首先训练专家策略模仿重新目标化的人类运动捕捉片段;其次将专家策略提炼为一个由本体感受编码器、残差向量量化(RVQ)码本和动作解码器组成的冻结架构;第三,在HMP保持冻结的情况下训练任务级策略通过选择离散码本条目来解决运动任务。我们在模拟中对速度跟踪、点目标导航和跌倒恢复速度跟踪进行评估,并在真实的宇树G1机器人上部署速度跟踪策略。提炼过程保留了专家的跟踪行为,而生成的HMP可作为不同下游运动策略的动作接口无需重新训练即可重复使用。学习到的HMP揭示了一种可解释的码本结构,其中活跃RVQ阶段的数量调节可用步态模式。我们还表明,与标准直通估计器相比,用旋转技巧训练码本可改善潜在组织并减少下游跌倒。
英文摘要
Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process. Motion imitation provides an alternative source of motor competence by training policies to track retargeted human motions, yet the resulting controllers remain reference trackers and are not directly usable as task policies. We propose a three-stage pipeline that turns motion-imitation skills into a reusable hybrid motion prior (HMP) for humanoid locomotion. First, an expert policy is trained to imitate retargeted human motion-capture clips. Second, the expert is distilled into a frozen architecture composed of a proprioceptive encoder, a residual vector-quantized (RVQ) codebook, and an action decoder. Third, task-level policies are trained to solve locomotion tasks by selecting discrete codebook entries while the HMP remains frozen. We evaluate the method on velocity tracking, point-goal navigation, and fall-recovery velocity tracking in simulation, and deploy the velocity-tracking policy on a real Unitree G1 robot. The distillation process preserves the tracking behavior of the expert, while the resulting HMP can be reused without retraining as the action interface for different downstream locomotion policies. The learned HMP reveals an interpretable codebook structure in which the number of active RVQ stages modulates the available gait patterns. We further show that training the codebook with the rotation trick improves latent organization and reduces downstream falls compared with a standard straight-through estimator.