arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

渐进式视觉规划的世界动作建模

World Action Modeling with Progressive Visual Planning

Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, Amir Bar

arXiv 2610.02508首次发表:更新:

发表机构

Shanghai Jiao Tong University; SII; Meta; Imperial College London(上海交通大学; SII; Meta; 伦敦帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProWAM通过联合预测动作和稀疏视觉子目标序列,提供渐进式视觉引导,实现高效长时程机器人控制,在模拟和真实世界基准上均取得最优性能。

AI 中文摘要

世界动作模型(WAMs)已成为机器人控制的一种有前景的范式,它通过从初始观察和指令中联合预测未来的视觉动态和动作来实现控制。然而,现有的WAMs在处理长时程预测时存在困难,因为生成密集的视频展开非常低效。近期一些WAMs通过仅预测单个未来帧而不生成完整视频来解决这一问题,但这种方法忽略了如何向目标推进。我们提出了ProWAM,一种渐进式世界动作模型,它联合预测动作和稀疏视觉子目标的顺序序列,为整个任务执行过程中的动作生成提供明确的视觉引导。这种设计具有自然的可扩展性,因为子目标预测可以从大规模无动作视频中学习,使得视频骨干网络能够将复杂的视觉规划从动作策略中卸载出来。为了实现高效的动作生成,ProWAM执行一次视频骨干网络的前向传播以缓存稀疏子目标特征,消除了迭代式完整视频生成,并在重新规划时仅需要轻量级的动作去噪。在广泛的评估中,ProWAM展现了卓越的分布外鲁棒性。在模拟基准测试中,它在LIBERO-Plus(85.8%)和随机化RoboTwin(75.7%)上取得了新的最先进结果,相比最强基线实现了高达+35.9%的相对提升。在RoboCasa365上,ProWAM实现了48.1%的成功率,在具有挑战性的Composite-Unseen分割中达到18.2%,总体排名第4。至关重要的是,在零样本真实世界实验中,ProWAM达到了70.0%的成功率,在新场景中比最强基线高出+15.0(从55.0%提升至70.0%,相对增益+27.3%)。这些结果证明了进度索引的视觉预见对于闭环控制的价值。我们的程序位于此https URL。

英文摘要

World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in https://sii-ferenas.github.io/ProWAM-page.

CommentsProject Page: https://sii-ferenas.github.io/ProWAM-page

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑