发表机构
University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示潜在世界模型中主导预测误差的组件未必是修复后最能改善决策的组件,通过分解误差并对比预言机修正,提出将表征、预测和规划作为独立阶段评估。
AI 中文摘要
主导潜在世界模型预测误差的组件未必是修复后最能改善动作选择的组件。我们通过比较相同物理起点下的动作序列,并将终点误差分解为候选池中心误差和动作相对响应误差来证明这一点。在四个模型家族和四个任务中,一个由每任务256个新起点和每起点300个共享候选组成的确认池显示,中心误差在16个模型-任务组合中的14个中主导均方误差(MSE)。然而,在其中的六个组合中,仅修正动作相对响应的预言机在物理秩相关和前30精英质量上优于仅修正中心的预言机,同时留下更多的潜在MSE(按家族修正的置信区间)。偏好因评估设置而异:一项单独的LeWorldModel(LeWM)研究执行预言机选择的动作,在PushT和具有渲染匹配目标的Reacher上偏向中心修复。匹配候选测试定位了排序损失:对于LeWM,在Reacher设置中编码实际终点将物理Spearman从0.464提升至0.975,在PushT上从0.193提升至0.631(每任务64个起点),而Cube的编码目标成本仍无信息量。一项72次运行的目标研究改善了所选响应的诊断,而增量闭环规划收益仍未得到确认。这些结果将误差幅度与预言机修正的决策效果分开,并激励将表征、预测和规划作为独立阶段进行评估。
英文摘要
The component that dominates a latent world model's prediction error need not be the one whose repair most improves action selection. We show this by comparing action sequences from identical physical starts and separating endpoint error into a candidate-pool center and action-relative responses. Across four model families and four tasks, a confirmation pool of 256 new starts per task and 300 shared candidates per start shows that center error dominates MSE in 14/16 model-task cells. Yet in six of these cells, an oracle that corrects only the action-relative responses yields better physical rank correlation and top-30 elite quality than one that corrects only the center, while leaving more latent MSE (family-wise corrected intervals). The preference differs across the evaluated settings: a separate LeWorldModel (LeWM) study that executes oracle-selected actions favors center repair on PushT and on Reacher with a render-matched goal. Matched-candidate tests localize ordering loss: for LeWM, encoding realized endpoints raises physical Spearman from 0.464 to 0.975 on that Reacher setting and from 0.193 to 0.631 on PushT (64 starts per task), while Cube's encoded-goal cost remains uninformative. A 72-run objective study improves selected response diagnostics, while incremental closed-loop planning gains remain unconfirmed. These results separate error magnitude from the decision effects of oracle correction and motivate evaluating representation, prediction, and planning as separate stages.
Comments40 pages, 7 figures