Kepler:用于ARC-AGI-3的可审计世界模型
Kepler: Auditable World Models for ARC-AGI-3
AI总结:
Kepler通过可执行世界模型和验证检查,在ARC-AGI-3上取得100.00 RHAE,并揭示了公开分数区分度有限的问题。
AI中文摘要:
ARC-AGI-3在交互式环境中评估智能体,这些环境的规则和目标必须通过观察来推断。我们提出了Kepler,一个开源框架,它将假设表示为可执行的世界模型,并通过回顾性转换检查和条件预测检查来验证它们。在一个固定的Claude Opus 5配置下,Kepler在所有25个公开游戏中获得了服务器验证的100.00 RHAE,且没有进行逐游戏模型选择或基于分数的条件重跑。在183个已完成的关卡中,有181个关卡的最终Opus尝试所使用的动作数不超过相应的中位数人类基线。保留的棋盘运行使用了8,256个环境动作,其中7,292个发生在计分关卡中。保留的本地提供商会话记录产生8.58亿个令牌,97.37%的缓存读取,以及按2026年9月1日API列表等价费率计算的777.72美元成本。我们还报告了三个评估失败:源代码泄漏导致无效的完美运行,智能体在控制条件下重建了移除的框架,以及自主修复掩盖了有缺陷的计划器。一个单游戏观察案例研究表明,动画帧包含任务相关信息,而这些信息在静态文本网格中缺失。在最终的Claude Opus 5和GPT-5.6 Sol棋盘上,50个游戏模型单元中有48个达到100。这些结果表明,仅公开集分数具有有限的区分价值,并激励首次尝试、成本条件和验证感知的报告。
英文摘要:
ARC-AGI-3 evaluates agents in interactive environments whose rules and objectives must be inferred from observation. We present Kepler, an open-source harness that represents hypotheses as executable world models and validates them through retrospective transition checks and conditional prediction checks. Under one frozen Claude Opus 5 configuration, Kepler obtained a server-verified 100.00 RHAE on all 25 public games, with no per-game model selection or score-conditioned reruns. On 181 of 183 completed levels, the final Opus attempt used no more actions than the corresponding median-human baseline. The retained board runs used 8,256 environment actions, of which 7,292 occurred in scored levels. Retained local provider-session records yield 858.0 million tokens, 97.37% cache reads, and a \$777.72 cost at September 1, 2026 API list-equivalent rates. We also report three evaluation failures: source-code leakage that produced an invalid perfect run, agents reconstructing a removed harness in a control condition, and autonomous repair masking a broken planner. A single-game observation case study showed that animation frames contained task-relevant information absent from settled text grids. Across the final Claude Opus 5 and GPT-5.6 Sol boards, 48 of 50 game-model cells reached 100. These results indicate that public-set score alone has limited discriminative value and motivate first-attempt, cost-conditioned, and verification-aware reporting.