发表机构
AGILabs(AGILabs实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型合成的代码世界模型,发现其虽预测准确率高,但在游戏中仍失败,存在验证与正确的差距,揭示危害规律,指出更多数据无法修复,相同机制在信念推理函数上重现,表明规划世界模型充分性应基于搜索分布或游戏衡量。
AI 中文摘要
大语言模型可将游戏规则合成为可执行代码,即代码世界模型(CWM),经典规划器会在其上进行搜索。通常当模型在采样轨迹上达到高转移准确率时就会被接受,但本文认为这对于规划来说是错误的充分性概念。研究表明:一是LLM合成的CWM能以100%转移准确率通过采样门且在规划器自身搜索分布上状态准确率≥98%,但在游戏中仍系统地失败,存在验证与正确的差距;二是危害遵循定量规律;三是更多数据无法修复失败,LLM合成表现为规则翻译而非推理;四是相同机制在不完全信息CWM的信念推理函数上重现,证明了覆盖界限。结果表明面向规划的世界模型的充分性应基于搜索分布或直接通过游戏来衡量,而非采样转移上的预测准确性。
英文摘要
Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches over. Such models are typically accepted when they reach high transition accuracy on sampled trajectories. We argue this is the wrong notion of adequacy for planning. We show four things. (1) An LLM-synthesized CWM can pass a sampling gate at 100% transition accuracy and be $\geq 98\%$ state-accurate on the planner's own search distribution, yet lose systematically at play, because the $<1\%$ it gets wrong is exactly the pivotal dynamics; the play cost of the omitted rule is $0.091$ (seed-clustered 95% CI $[0.065,0.117]$, $n=4800$). We call this the verified-vs-correct gap, and confirm it end-to-end through the synthesis pipeline. (2) The harm follows a quantitative law, $\mathrm{danger}=\mathrm{play\_cost}\times(1-\mathrm{rarity})^N$, whose $(1-\mathrm{rarity})^N$ gate-miss factor is proven exact and whose play cost is empirically bounded. (3) The failure is not repaired by more data: LLM synthesis behaves as rule translation, not rule inference, and did not infer the omitted rule across models (GPT-5.x) and data regimes (including DAgger and targeted examples). (4) The same mechanism recurs on the belief-inference function of imperfect-information CWMs: we prove a coverage bound (a size-$N$ gate is identifying when $N\gtrsim b^{d_{\max}}$), explaining why shallow games such as Kuhn poker show no gap, and hand-construct Beacon, a verified-but-wrong inference function that passes the gate yet loses every game. These results suggest adequacy for planning-oriented world models should be measured on the search distribution or by play directly, not by prediction accuracy on sampled transitions.
Comments41 pages, 4 figures. Code and reproduction log: https://github.com/JaviMaligno/code-world-models