发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SimForcing框架,通过潜在运动蒸馏将仿真运动先验迁移至真实域机器人世界模型,并引入带条件丢弃的多块仿真条件化与无分类器引导,在Bridge等数据集上取得最优视频生成指标,并提升下游策略学习性能。
AI 中文摘要
基于动作条件的机器人世界模型必须精确响应机器人轨迹,同时保持真实的视觉动态,然而从异构机器人视频中同时学习这两方面仍然具有挑战性。仿真提供了结构化的运动监督,但外观差异阻碍了直接迁移,且不准确的仿真预测可能误导真实视频生成。我们提出了SimForcing,一种仿真引导的框架,将仿真既用作可迁移运动知识的来源,又用作可控的预测参考。首先,我们通过潜在运动蒸馏从仿真教师中迁移运动知识,在潜在空间中对齐时间变化以内化运动先验,同时减轻外观差异的影响。其次,我们引入了带条件丢弃的多块仿真条件化,以利用预测的仿真轨迹而不过度依赖其准确性。我们的仿真条件化无分类器引导方案通过平衡基于内化运动知识的预测与额外由仿真潜在变量引导的预测,统一了这两种思想。联合训练的学生模型同时生成仿真条件和真实域视频,推理时无需额外的世界模型。在Bridge上,SimForcing在比较方法中取得了最佳的PSNR、SSIM、LPIPS和FVD,且无需外部具身预训练。在InternData-A1上的评估进一步支持了其跨机器人数据集的适用性。此外,使用我们训练的世界模型初始化视觉-语言-动作模型提高了LIBERO的成功率,表明其对于下游策略学习具有实用性。\url{这个https URL}
英文摘要
Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. \url{https://github.com/Wang-Xiaodong1899/SimForcing}
CommentsCode: https://github.com/Wang-Xiaodong1899/SimForcing