发表机构
Massachusetts Institute of Technology; Harvard University(麻省理工学院; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对预测性世界模型生成展开慢的问题,提出基于漂移生成模型的DriftWorld,训练时学习动作条件漂移以快速生成未来帧,在机器人操作基准测试中实现快速准确决策,还能作离线模拟器,性能优于基于扩散的基线。
AI 中文摘要
预测性世界模型能让机器人通过想象行动结果来进行规划,但其对控制的价值取决于能否快速生成多个展开。这给基于扩散的世界模型带来瓶颈:多步采样使每个展开成本高昂,限制了推理时的大规模动作搜索。我们引入DriftWorld,一种基于漂移生成模型的动作条件世界模型。它在训练时学习动作条件漂移,而非在推理时迭代去噪,能在单次前向传播中以30+帧每秒的速度从当前观察和候选动作序列生成未来帧,比基于扩散的基线平均快17倍。我们在标准视觉机器人操作基准上评估DriftWorld,它生成的展开既准确又快速,在推理时间远少于基于扩散的世界模型基线的情况下实现了最优决策性能。此外,DriftWorld还可作为离线模拟器对现实世界机器人策略进行排序,基于展开的分数与地面真值的相关性高达0.99。这些结果表明漂移模型非常适合机器人世界建模,快速、高质量的想象能直接支持规划和策略评估。
英文摘要
Predictive world models enable robots to simulate the visual outcomes of their actions, but state-of-the-art diffusion-based models remain costly because generating each rollout requires multi-step iterative denoising. We introduce DriftWorld, an action-conditioned world model based on drifting generative models. DriftWorld learns a conditional drift during training, enabling it to generate future observations for a given action sequence in a single forward pass during inference. Across Bridge-V2, RT-1, Language Table, Push-T, and Robomimic, DriftWorld runs at over 40 fps and is 12+ times faster than diffusion-based baselines, while matching or improving their visual generation quality. This makes DriftWorld an efficient world model for robot simulation and further enables downstream applications including inference-time action search and offline policy evaluation.
CommentsWebsite at https://susie-lu.github.io/driftworld/