AI 中文总结
该研究提出Twin系统,通过测试时数字孪生构建可执行世界模型,在ARC-AGI-3等未知游戏中通关率达97.8%,效率优于人类,核心是通过模拟交互和反例修复推断游戏规则与目标。
AI 中文摘要
我们提出了一种测试时世界模型推理(Test-time World-model Inference,Twin)系统,其中前沿编码智能体会编写可执行的世界模型,用于完成诸如ARC-AGI-3游戏之类的持续学习任务。传统方法会手动设计这类模型,每个任务对应一种定制设计。每个游戏都隐藏了其规则和目标,而我们的系统仅通过模拟和交互来构建这些规则与目标。它对网格游戏的归纳先验足够强大,能够在几乎所有关卡中恢复游戏的真实状态转移和目标。孪生世界模型中会进行重放验证:在程序复现每一个先前观测到的游戏状态转移之前, harness(框架)会强制不执行任何动作。世界模型预测与实际动作结果之间的每一个不匹配,都会成为用于修复世界模型的反例。Twin在183个关卡中通关了179个(97.8%),且在179个已通关关卡中的158个(88.3%)上比人类表现更高效;在其通关的关卡中,有156个(87.2%)是在获得任何奖励之前就推断出了目标,其余关卡则通过搜索自动发现目标。该基准针对人类首次玩每个游戏的表现,从0到100分对通关情况和动作效率进行评分:直接使用基础模型仅得7.8分;现成的harness可将其提升至61.1分,而我们的孪生世界模型可将同一基础模型提升至93.3分,在25个游戏中通关了23个。构建可用的世界模型比预期更简单,而更难的问题是推断出正确的目标。
英文摘要
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom design per task. Each game hides its rules and goal, and our system constructs them from simulation and interaction alone. Its inductive prior over grid games is strong enough to recover the true transitions of the game and the goal on nearly all levels. Replay validation happens in a twin world model. The harness enforces that an action is not made until the program reproduces every previous observed game transition. Each mismatch between a world model prediction and the actual action result becomes a counterexample that is used to repair the world model. Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%). The system infers the goal before any reward on 156 of the levels it clears (87.2%), and in the remaining levels automatically discovers the goal by search. The benchmark scores completion and action efficiency, between 0 and 100, against humans playing each game for the first time. Played directly, the base model scores only 7.8%; an off-the-shelf harness increases it to 61.1%, whereas our twin world model increases the same base model to 93.3%, clearing 23 out of 25 games. Building a usable world model is simpler than anticipated, whereas the harder problem is inferring the right goal.
CommentsProject website with action-by-action replays of all 25 runs: https://arc-agi-3-twin.vercel.app/ Code: AGI-3" target="_blank" rel="noopener">https://github.com/Alexyskoutnev/TWIN-ARC-AGI-3