arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

目标是瓶颈:潜在世界模型编码了其规划器无法使用的内容

The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

Joyjeet Singh

arXiv 2608.12959首次发表:更新:

AI 中文总结

该研究发现潜在世界模型的规划瓶颈在于规划器目标而非预测器,通过替换目标无需额外训练即可大幅提升长视野规划性能,揭示模型编码了规划器未利用的可达性信息。

AI 中文摘要

潜在世界模型的评判标准是其预测效果,因此当长视野规划失败时,自然会认为是预测器性能下降。我们在TwoRoom环境中复现LeWorldModel时发现,约束规划的关键因素是规划器的目标而非预测器。预测器并非性能瓶颈:其想象的75个环境步后的状态,错误程度仍仅为假设世界冻结时的0.189,而规划器的想象步长从未超过25。问题出在目标上:交叉熵方法规划最小化潜在空间距离,该距离在r=0.426时与真实距离相关,在约80个竞技场单位处达到饱和,超过120后反而下降,即远离目标会降低代价。信息全程存在:脊线探测从冻结嵌入中恢复位置的R²值为0.9922。该缺陷是方法本身的问题,而非复现问题,它存在于作者发布的权重中,且在四个检查点上,长视野成功排名与度量质量正相关、与预测准确率负相关。仅替换目标(无需重新训练、无需GPU),可将偏移100时的目标达成率从26.0%提升至98.0%,达到偏移25时的98.0%,且在预算仅为原来三分之一时达到92.0%,使规划不再依赖视野长度。最优代价并非最准确:仅从帧分离学习的头模块,预测空间距离的能力弱于位置探测(r=0.819,对比0.9897),但规划效果更好,穿越环境分隔墙的代价高24%,而平方潜在距离代价仅高4%,该模块学习到的是可达性而非邻近性。

英文摘要

Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as wrong as assuming the world froze, while the planner never imagines beyond twenty-five. The objective is. Cross-entropy-method planning minimises squared latent distance, which tracks true distance at r = 0.426, saturates by about eighty arena units and decreases beyond a hundred and twenty, so moving away from the goal can lower the cost. The information is present throughout: a ridge probe recovers position from the frozen embedding at R^2 0.9922. The pathology is the method's, not one reimplementation's. It is present in the authors' released weights, and across four checkpoints long-horizon success rank-orders exactly with metric quality and inversely with prediction accuracy. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0%, equals the 98.0% at offset 25, and reaches 92.0% under a third of the budget: planning stops depending on the horizon. The best cost is not the most accurate. A head learned from frame separation alone predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better, charging 24% more to cross the environment's dividing wall where squared latent distance charges 4% less. It has learned reachability, not proximity.

CommentsFollow-up to arXiv:2608.10145. All experiments run on a laptop CPU; no model was trained or fine-tuned. Code, checkpoints and every measurement: github.com/joyjeet-singh/tinylab

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑