世界时间计算与经过验证的代码世界模型
World-Time Compute with Verified Code World Models
浏览论文内容
中文总结 AI 辅助
本文提出世界时间计算,通过代码世界模型生成精确标记轨迹进行训练,提升LLM对未见世界的泛化能力,在0.5B模型上提升29个百分点,并验证了标签精确性驱动收益。
中文摘要 AI 辅助
大语言模型只有在看到大量真实且带标签的示例后才能在某个领域实现泛化,而大多数领域缺乏这样的数据。我们研究了一种廉价制造这种泛化能力的方法。当一个领域的动态可以用代码表示时,一个模板可以实例化为许多世界模型:这些模型是在符号状态上可执行、可验证的程序,每个模型都是精确标记轨迹的取之不尽的来源。在这些世界中,通过对许多这样的世界中的轨迹进行微调(我们称之为世界时间计算,即测试时计算的训练时对应物),可以提升对从未训练过的保留世界(合成世界族)的泛化能力。收益在能力最稀缺的地方最大:在0.5B规模下提升29个百分点;最大模型的提升在噪声范围内,与饱和效应一致。标签可以被信任,因为世界是经过验证的代码:合成后检查的动态在20步轨迹上是精确的,并且对10倍分布外探测的回答完全正确(100%),而逐步的LLM和MLP预测器则累积误差并崩溃。与域随机化不同,每个世界都是独立编写和验证的;一个带损坏标签的对照组表明,标签的精确性而非任务多样性驱动了收益。在真实基准(ARC-AGI网格、List Functions、CLRS)上,同样的杠杆与逐世界测试时训练一样有效。在List Functions上,更难的跨世界形式成立:一个在128个不相交世界上训练的适配器在保留世界上达到40%,而损坏标签对照组为6%(提升34个百分点,置信区间[29, 39])。这种收益是一种饱和的规律性而非定律:对少步推理和小型/弱模型最大,对长链、感知引发的任务和饱和任务逐渐减弱;没有共享技能时跨任务迁移较弱。世界由OpenWorld(一个零依赖框架,见配套论文)编写和服务。范围:符号状态;像素原生领域仍是学习模型的领域。所有代码、配方和本手稿都从一个仓库重新生成。
英文摘要
LLMs generalize across a domain only after seeing many real, labeled examples, which most domains lack. We study a way to manufacture it cheaply. When a domain's dynamics can be written as code, one template instantiates into many world models: executable, verifiable programs over symbolic state, each an inexhaustible source of exactly-labeled trajectories. Fine-tuning an LLM on trajectories through many such worlds, which we call world-time compute, a training-time analogue of test-time compute, lifts generalization to held-out worlds it never trained on (synthesized world families). Gains are largest where capability is scarcest: +29 points at 0.5B; the largest model's lift is within noise, consistent with saturation. Labels can be trusted because the worlds are verified code: synthesized-then-checked dynamics are exact over 20-step rollouts and answer 10x out-of-distribution probes exactly (100%), whereas per-step LLM and MLP predictors compound error and collapse. Unlike domain randomization, each world is independently authored and verified; a corrupted-label control shows label exactness, not task variety, drives the gains. On real benchmarks (ARC-AGI grids, List Functions, CLRS) the same lever holds as per-world test-time training. On List Functions the harder cross-world form holds: one adapter trained on 128 disjoint worlds reaches 40% on held-out worlds versus 6% for a corrupted-label control (+34 points, CI [29, 39]). The gain is a saturating regularity, not a law: largest for few-step reasoning and small/weak models, fading for long chains, perception-induced tasks, and saturated tasks; cross-task transfer is weak without shared skill. Worlds are authored and served by OpenWorld, a zero-dependency framework (companion paper). Scope: symbolic state; pixel-native domains remain territory of learned models. All code, recipes, and this manuscript regenerate from one repository.
发表机构
- Quome, Inc.(Quome公司)
- University of Washington(华盛顿大学)
- Rafter Labs, Inc.(Rafter实验室公司)
- OWASP(OWASP(开放Web应用程序安全项目))
- Health Stream Analytics, LLC(Health Stream Analytics有限责任公司)
- IEEE Computer Society(IEEE计算机学会)
- Fusen World LLC(Fusen World有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。