arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在世界模型的规划极限

The Planning Limits of Latent World Models

Ali Alrasheed, Basim Azam, Naveed Akhtar

arXiv 2609.39235首次发表:更新:

发表机构

The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示了潜在世界模型的规划极限:其预测仅在目标位于想象轨迹附近时可靠,且此范围无法通过扩大模型或延长训练扩展,但在此范围内可显著提升VLA策略成功率。

AI 中文摘要

世界模型为帮助机器人理解物理世界如何演化并通过想象规划复杂行为提供了一种有前景的方法。然而,现有研究主要展示了这些模型能够完成什么,而对其预测在规划中何时仍然有用以及在何处失效的问题尚不明确。我们利用基于五个冻结的自监督视觉骨干网络构建的动作条件预测器来研究这一问题,这些骨干网络包括V-JEPA 2、V-JEPA 2.1、VideoMAEv2、VideoPrism和DINOv2。我们使用冻结的骨干网络来测试旨在跨环境迁移的表征。我们在多样化的Meta-World操作任务以及BridgeData V2的真实机器人交互上评估了这些模型。我们发现,只有当目标位于规划期间模型所想象的轨迹之内或略微超出该轨迹时,世界模型才能可靠地指导动作选择。在五步展开(即预测器训练时所用的长度)的情况下,世界模型仅能对超前五到十个控制步骤的目标可靠地对动作进行排序,而任务目标则位于十六到五十三个步骤之外。无论是将预测器扩大八十一倍,还是进行更长展开的训练,都无法扩展这一范围;编码器同时影响排序范围和闭环成功率,其中V-JEPA 2.1的表现最为稳定。更根本的是,这一限制在完美预测下依然存在:使用真实模拟器,当目标从五步展开超前五步移动到超前二十步时,成功率从92%下降到41%。因此,规划需要更长的想象轨迹或更近的子目标。对于遥远的目标,纯想象在23%的回合中成功,带反馈的规划(MPC)将成功率提升至30%,想象至目标处则提升至47%,而使用邻近的专家子目标则提升至76%。在其可规划的范围内使用世界模型,还可以改进视觉-语言-动作(VLA)策略:从VLA提出的八个动作中选择一个,可将16个不同任务上的成功率从65%提升至77%。

英文摘要

World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet existing studies mainly demonstrate what these models can accomplish, leaving unclear when their predictions remain useful for planning and where they fail. We study this question using action-conditioned predictors built on five frozen self-supervised visual backbones: V-JEPA 2, V-JEPA 2.1, VideoMAEv2, VideoPrism, and DINOv2. We use frozen backbones to test representations intended to transfer across environments. We evaluate these models on diverse Meta-World manipulation tasks and real-robot interactions from BridgeData V2. We find that a world model guides action selection reliably only when the goal lies within, or slightly beyond, the trajectory it imagines during planning. With five-step rollouts, the length the predictor was trained on, the world model ranks actions reliably only for targets five to ten control steps ahead, whereas task goals lie 16 to 53 steps away. Neither an 81-fold larger predictor nor longer-rollout training extends this range; the encoder affects both range and closed-loop success, with V-JEPA 2.1 performing most consistently. More fundamentally, the limit persists under perfect prediction: using the real simulator, success falls from 92% to 41% as the target moves from five to twenty steps ahead of a five-step rollout. Planning therefore requires either longer imagined trajectories or closer subgoals. For distant goals, pure imagination succeeds in 23% of episodes, planning with feedback (MPC) raises success to 30%, imagining as far as the goal to 47%, and nearby expert subgoals to 76%. Used within its plannable range, a world model can also improve a vision-language-action (VLA) policy: choosing among eight actions the VLA proposes raises its success from 65% to 77% across 16 different tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑