评估协议决定结果:在TwoRoom上独立复现LeWorldModel
The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom
浏览论文内容
中文总结 AI 辅助
本研究独立复现LeWorldModel在TwoRoom上的结果,发现评估协议(如目标偏移量)会显著影响性能,还揭示单步预测准确率无法预测长程规划成功、批量归一化层会夸大验证损失等关键结论。
中文摘要 AI 辅助
LeWorldModel利用预测损失和单一抗坍缩正则化器训练潜在世界模型,并报告在其最简单的诊断环境TwoRoom上达到约87%的目标完成率。我们通过独立重新实现,在价值约25美元的租用计算资源上完成复现,所有评估均在一台笔记本CPU上进行。在仓库的评估目标偏移量下,我们达到94.0%的完成率;而在相同情节下,使用我们的协议评估作者自行发布的检查点,其完成率为84.0%。我们直接复现了报告的表示结果:位置探针的皮尔逊相关系数r=0.9988,与报告的0.996相符。要达到这一结果需要四个决定结果的约定,且这些约定未出现在任何发布的配置文件中:帧跳过块内的密集动作收集、编程设置的动作编码器宽度、ImageNet像素归一化以及动作Z分数标准化。仅遵循发布配置的复现者会得到一个预测器无法收敛的模型。发布的材料本身对评估协议存在争议:论文附录和仓库配置指定了不同的目标偏移量和步数预算;使用作者自身权重时,这两种设置分别得到14.0%和84.0%的完成率,且仅配置的值能复现报告的数字。在50个相同情节中,仅改变目标的构建方式,该检查点的完成率就从84.0%降至8.0%。有两个发现具有普遍性:一是单步预测准确率无法预测长程规划成功——在跨越7倍预测误差范围的三个检查点(包括作者自身的)中,其能单调排序短程成功,但完全无法排序长程成功;二是批量归一化层使我们报告的验证损失最多膨胀300倍,掩盖了全程平稳的训练损失。
英文摘要
LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment. We reproduce that result by independent reimplementation on roughly $25 of rented compute, with all evaluation on one laptop CPU. We reach 94.0% at the repository's evaluation goal offset, against 84.0% for the authors' own released checkpoint measured under our protocol on identical episodes, and we reproduce the reported representation result directly (position probe Pearson r = 0.9988 against a reported 0.996). Reaching that point required four conventions that determine the outcome and appear in no released configuration file: dense action gathering across a frameskip block, a programmatically-set action-encoder width, ImageNet pixel normalisation, and action z-scoring. A reproducer following the released configurations alone obtains a model whose predictor cannot converge. The evaluation protocol is itself contested by the released material. The paper's appendix and the repository's configuration specify different goal offsets and step budgets; on the authors' own weights these yield 14.0% and 84.0%, and only the configuration's values reproduce the reported figure. On fifty identical episodes, changing nothing but how the goal is constructed moves that checkpoint from 84.0% to 8.0%. Two findings generalise. One-step prediction accuracy does not predict long-horizon planning success: across three checkpoints spanning a sevenfold range in prediction error, including the authors' own, it orders short-horizon success monotonically and fails to order long-horizon success at all. And a batch normalisation layer inflated our reported validation loss by up to a factor of 300, concealing a training loss that was flat throughout.