发表机构
NTU; BAAI; PKU; NJU(南洋理工大学; 北京智源人工智能研究院; 北京大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RoboDreamer提出两阶段教师-学生框架,结合下一观测一致性与随机时间掩蔽,利用Mamba骨干网络实现人形机器人在不完美感知下的鲁棒运动跟踪,并在仿真和真实机器人上验证。
AI 中文摘要
人形机器人运动需要控制策略在不完美感知下保持稳定,同时利用时间上下文实现一致的运动。我们提出RoboDreamer,一个两阶段的教师-学生框架,结合了下一观测一致性与随机连续时间掩蔽。首先在干净观测上训练教师模型,然后在掩蔽近期观测下蒸馏学生模型,鼓励策略从历史中推断缺失的当前信息。在推理时,复用相同的掩蔽接口进行隐式闭环动作细化和可选的多步动作分块。Mamba用作时间骨干网络,匹配的消融实验表明,掩蔽/蒸馏提供了大部分性能提升,而Mamba在实时延迟下贡献了额外的跟踪改进。在IsaacLab、MuJoCo和Unitree G1上的实验表明,在观测掩蔽下具有鲁棒的运动跟踪,并成功实现了真实世界部署。
英文摘要
Humanoid locomotion requires control policies that remain stable under imperfect sensing while exploiting temporal context for consistent motion. We present RoboDreamer, a two-stage teacher--student framework that combines next-observation consistency with randomized continuous temporal masking. A teacher is first trained on clean observations, and a student is then distilled under masked recent observations, encouraging the policy to infer missing current information from history. At inference, the same masking interface is reused for implicit closed-loop action refinement and optional multi-step action chunking. Mamba is used as the temporal backbone, while matched ablations show that masking/distillation provides a substantial part of the gain and Mamba contributes additional tracking improvements with real-time latency. Experiments in IsaacLab, MuJoCo, and on a Unitree G1 demonstrate robust motion tracking under observation masking and successful real-world deployment.