arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36413cs.RO

从无限中取一:将预训练世界模型的未来转化为机器人动作

One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions

  • University of California, Berkeley(加州大学伯克利分校)
  • Southern University of Science and Technology(南方科技大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Bang Du, Yichen Xie, Shuqi Zhao, Yuxin Chen, Menglin Wu, Masayoshi Tomizuka

中文总结 AI 辅助

本文提出RoboActualizer,利用冻结世界模型,以6000万参数的小模型实现任务条件化的未来选择与动作生成,性能优异且延迟低至39毫秒。

中文摘要 AI 辅助

预训练的视频世界模型为场景提供了许多可能的未来,但机器人必须实现精确的任务条件化未来。为了将世界模型转化为可执行的机器人策略,现有方法使用大规模机器人数据和计算资源对重型世界模型骨干进行微调。挑战这一现状,我们认为昂贵的部分已在世界模型预训练中支付,因为视频世界模型的表示空间展现了多样化的潜在未来。在这种情况下,剩下的就是选择完成任务的未来并读出实现它的动作。我们将此任务形式化为“实际化”,即在冻结世界模型提供的先验之上学习任务条件化的选择和实现。这可以通过一个微小的实际化器模型来解决。我们实现了RoboActualizer,其参数仅为6000万,位于冻结的世界模型编码器之上。该实际化器由两个轻量级DiT专家组成,通过流匹配联合预测未来的潜在表示和动作。该模型可以完全在单个GPU上训练,峰值内存为32 GB。与现有的WAM和VLA相比,可训练参数最多减少100倍,RoboActualizer在仿真基准(包括LIBERO、LIBERO-Plus、RoboTwin 2.0)以及两个真实平台上的五项任务中达到了出色的性能,低延迟为39毫秒,可实现实时控制。

英文摘要

A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.

↑