AI 中文总结
XPACE提出统一具身世界模型,通过共享视频骨干联合预测动作与未来视频,利用异构经验训练并自生成恢复数据,提升人形机器人任务完成与技能迁移能力。
AI 中文摘要
通用机器人需要利用多样化的经验、选择动作,并预测这些动作将如何改变世界。我们提出XPACE,一个统一的具身世界模型,它既充当世界动作模型,联合预测可执行的机器人动作和未来视频,又充当世界模拟器,预测指定动作的视觉后果。我们的关键洞见在于,视频预测既能将异构经验与动作学习联系起来,又能为策略改进生成新经验。通过在策略和模拟器之间共享视频骨干网络,我们利用无动作标签的视频学习视觉动态,并利用带动作标签的人类和机器人演示联合学习视频和动作预测。基于此架构,一种从粗到细的训练课程逐步强调机器人控制,同时保留人类经验,使策略能够学习超出机器人演示范围的行为。除了从记录的经验中学习外,XPACE还利用其模拟器为策略生成额外的恢复监督。具体而言,我们使模拟器适应其自身生成的上下文,围绕专家演示合成偏差-恢复轨迹,并在过滤后的恢复示例上微调策略。在XPENG的IRON人形机器人上的实验表明,异构训练提高了鲁棒性,并能够将人类观察到的技能迁移到机器人演示中不存在的任务上,而模型自身模拟器生成的恢复数据进一步提高了真实世界的任务完成率。这些结果共同证明了世界与动作联合建模如何将异构经验学习与模拟驱动的策略自我改进联系起来。
英文摘要
A general-purpose robot needs to draw on diverse experience, choose actions, and anticipate how those actions will change the world. We introduce XPACE, a unified embodied world model that serves as both a world action model, jointly predicting executable robot actions and future video, and a world simulator, predicting the visual consequences of prescribed actions. Our key insight is that video prediction can both connect heterogeneous experience to action learning and generate new experience for policy improvement. With a shared video backbone between the policy and simulator, we use action-unlabeled video to learn visual dynamics and action-labeled human and robot demonstrations to jointly learn video and action prediction. Building on this architecture, a coarse-to-fine training curriculum progressively emphasizes robot control while retaining human experience, allowing the policy to learn behaviors beyond those covered by robot demonstrations. Beyond learning from recorded experience, XPACE uses its simulator to create additional recovery supervision for the policy. Specifically, we adapt the simulator to its own generated context, synthesize deviation-recovery trajectories around expert demonstrations, and fine-tune the policy on filtered recovery examples. Experiments on XPENG's IRON humanoid robot show that heterogeneous training improves robustness and enables transfer of human-observed skills to tasks absent from robot demonstrations, while recovery data generated by the model's own simulator further improves real-world task completion. Together, these results demonstrate how joint world and action modeling connects learning from heterogeneous experience with simulation-driven policy self-improvement.