目标放近,方能致远:冻结的世界模型比你想的更能规划
Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出锚定规划,通过检索与当前和目标观察相似的记录片段并瞄准中间目标,使冻结世界模型在无需训练的情况下超越最终目标评分,显著提升长程控制性能。
中文摘要 AI 辅助
基于视觉世界模型的规划器通常根据每个预测结果与编码的目标图像之间的距离来评分。我们表明,即使具有精确的动力学和全局最优的短视搜索,这一目标也可能限制控制:达到目标可能需要最初远离目标的动作。使用冻结的LeWM模型,中间目标在Cube、PushT、Reacher和TwoRoom任务上显著改善了动作合成和记录动作排名。学习到的目标和从观察经验中抽取的目标都能产生这些收益。我们引入了锚定规划(Anchored Planning),它检索一个记录片段,其起点和终点分别类似于当前观察和目标观察,然后瞄准该片段起点之后不久的一个观察。冻结模型从当前状态对朝向该目标的动作进行评分。无需额外训练,在长程评估的每项任务中,朝向观察目标的规划都优于发布的LeWM规划器。额外的最终目标搜索未能达到同样的收益。较低的后继预测误差不一定转化为更好的控制。成功还取决于目标放置的远近,以及随着执行推进而缩小检索跨度。仅改变目标,就能让相同的冻结模型和规划器达到最终目标评分所遗漏的目标。
英文摘要
Latent world models plan toward goal images with a frozen pretrained predictor, without task rewards or extra trained heads. However, their planners struggle with long-range goals, and prior work addresses this by training extra components such as value functions or subgoal models. We show that the planning target itself can cause this failure: even with exact dynamics and globally optimal short-horizon search, scoring predictions by their distance to the final goal rejects the first steps of a route that initially moves away from the goal. Building on this insight, we propose Anchored Planning (AP), a training-free method that reuses the world model's own offline trajectories. AP retrieves a segment that leads from the current observation toward the goal and aims the frozen planner at an observation shortly after the segment's start. Across four diverse tasks, AP substantially improves frozen LeWM planners for both action synthesis and action ranking, and it outperforms both additional final-goal search and the LeWM planner on long-range goals.