Pelican-Sim 1.0:面向具身智能的通用世界模型模拟器
Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
- Beijing Innovation Center of Humanoid Robotics (X-Humanoid)(北京人形机器人创新中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出通用世界模型模拟器Pelican-Sim 1.0,通过统一动作表示、动作-视觉注入、稀疏MoE和高效滚动生成,在约百万轨迹上训练,显著提升动作可控性和视频质量,并成功支持多项下游任务。
AI中文摘要:
在本技术报告中,我们提出了Pelican-Sim 1.0,一个面向具身智能的通用世界模型模拟器,它能够根据视觉上下文和机器人动作预测未来观测,以支持下游学习和决策。该模型包含四个关键设计特性:(1)统一动作表示:一个28维动作值空间覆盖了大多数主流具身形态,使得一个模型在异构设备上保持有效。(2)动作-视觉注入:基于URDF和相机渲染的动作视频将动作与像素联系起来,在不同具身形态、场景和任务中显著提高了可控性(相比替代融合基线,PSNR提升0.904)。(3)稀疏混合专家(MoE):稀疏MoE层增加了对异构动力学的容量,并吸收了动作模态,同时减少了模态间冲突(相比稠密主干,FVD降低6.530)。(4)高效滚动生成:因果适应和少步蒸馏产生了一个四步自回归模拟器,相比35步模型实现了5.67倍的加速。得益于这些设计,我们在约一百万条真实世界和模拟轨迹上进行了训练,并在动作可控性和视频质量方面获得了大幅提升:在AgiBotWorld Beta上,PSNR相比最强评估基线提升了4.636,在RoboMIND上提升了2.080,在RoboTwin上提升了10.343,且适配后的EWMBench DYN分数在RoboTwin上提升了0.426。基于此,RoboTwin上的四个下游应用取得了成功:每个任务在50个演示基础上添加500条生成轨迹,将策略成功率从70%提升至93%;策略评估在五个检查点上达到了0.994的皮尔逊相关系数;相对成功率提升在动作选择上达到47.7%,在策略改进上达到20.3%。在轨迹、场景、物体、具身形态和视点变化上的定性泛化突显了其作为通用世界模型模拟器的潜力。
英文摘要:
In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-rendered action videos bridge actions and pixels, giving markedly better controllability across embodiments, scenes, and tasks (PSNR +0.904 over alternative fusion baselines). (3) Sparse mixture-of-experts (MoE): sparse MoE layers add capacity for heterogeneous dynamics and absorb the action modality while reducing inter-modality conflict (FVD -6.530 vs. the dense backbone). (4) Efficient rollout generation: causal adaptation and few-step distillation yield a four-step autoregressive simulator, achieving a 5.67-fold speedup over the 35-step model. Benefiting from these designs, we train on approximately one million real-world and simulated trajectories and obtain large gains in action controllability and video quality: PSNR improves over the strongest evaluated baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin, with the adapted EWMBench DYN score up 0.426 on RoboTwin. Relying on this, four downstream applications on RoboTwin succeed: 500 generated trajectories added to 50 demonstrations per task raise policy success from 70% to 93%; policy evaluation reaches a Pearson correlation of 0.994 across five checkpoints; and relative success gains reach 47.7% for action selection and 20.3% for policy improvement. Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights its potential as a general-purpose world model simulator.