DreamMimic:通过世界模型学习视觉运动全身移动操作
DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model
浏览论文内容
中文总结 AI 辅助
DreamMimic 框架通过世界模型辅助蒸馏,将特权教师策略提炼为基于视觉的人形机器人控制器,引入 PCG 平衡引导与探索,在 OMOMO 和 BEHAVE 上提升了视觉移动操作性能。
中文摘要 AI 辅助
基于视觉的人形机器人全身移动操作极具挑战性,原因在于部分可观测性、接触丰富的动力学,以及从高维视觉输入中学习长 horizon 行为的难度。我们提出 DreamMimic 框架,该框架通过世界模型辅助蒸馏,将特权教师策略提炼为基于视觉的人形机器人控制器。我们未采用 Dreamer 风格的 RSSM 进行规划,而是将其重新用于学习预测性潜态动力学,该动力学同时作为表示空间和动作条件的多步监督信号,同时向学生策略暴露紧凑的预测特征以减少长期漂移。除了针对本体感受和视觉观测的标准重建目标外,我们还添加了辅助预测头,用于特权状态、接触、物体状态和奖励估计。这些头提供了与智能体-物体交互和任务进展相关的额外监督,鼓励潜态表示保留对接触丰富的移动操作有用的信号。我们进一步引入性能条件引导(PCG),一种奖励驱动的自适应蒸馏调度,该调度计算教师和学生的性能分数以动态平衡引导与探索。PCG 可防止在具有挑战性的视觉设置中过早的教师退火和过度的教师干扰。在 OMOMO 和 BEHAVE 上的实验表明,与强大的基于视觉的基线相比,跟踪式移动操作性能得到提升,且部署时未向学生暴露在线特权交互状态。定性模拟进一步研究了形态和模拟器变化。这些结果表明,世界模型可提供一种有用的机制,用于稳定接触丰富的人形机器人行为中的视觉策略蒸馏。
英文摘要
Vision-based whole-body loco-manipulation on humanoid robots is challenging due to partial observability, contact-rich dynamics, and the difficulty of learning long-horizon behaviors from high-dimensional visual inputs. We present \href{https://github.com/DreamMimic/DreamMimic}{DreamMimic}, a framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation. Instead of using a Dreamer-style RSSM for planning, we repurpose it to learn predictive latent dynamics that serve as both a representation space and an action-conditioned multi-step supervision signal, while exposing compact predictive features to the student policy to reduce long-term drift. Beyond standard reconstruction objectives for proprioceptive and visual observations, we add auxiliary prediction heads for privileged state, contact, object state, and reward estimation. These heads provide additional supervision related to agent--object interaction and task progress, encouraging the latent representation to retain signals that are useful for contact-rich loco-manipulation. We further introduce Performance-Conditioned Guidance (PCG), a reward-driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance guidance and exploration. PCG prevents both premature teacher annealing and excessive teacher interference in challenging visual settings. Experiments on OMOMO and BEHAVE show improved tracking-based loco-manipulation performance over strong vision-based baselines, without exposing online privileged interaction states to the student at deployment. Qualitative simulations further examine morphology and simulator changes. These results suggest that world models can provide a useful mechanism for stabilizing visual policy distillation in contact-rich humanoid behaviors.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。