arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tycho:面向ARC-AGI-3的基于可编程世界模型的主动抽象

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

Jens Lehmann, Andrei Aioanei, Sahar Vahdati

arXiv 2607.28287首次发表:更新:

发表机构

Dresden University of Technology; Amazon; TIB – Leibniz Information Centre for Science and Technology; Leibniz University of Hannover(德累斯顿工业大学; 亚马逊; 莱布尼茨科学与技术信息中心; 汉诺威莱布尼茨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对ARC-AGI-3的交互式抽象问题,提出编码智能体系统Tycho,通过结构化观测构建游戏模型,选定策略下GPT-5.6 Sol与Opus 5可达成100%相对人类行动效率并完成全部关卡。

AI 中文摘要

ARC-AGI-3将抽象问题转化为技能获取的交互式问题,玩家必须推断陌生游戏的规则、隐藏状态与目标,同时保持行动效率,因为每一步都至关重要。我们将这些环境形式化为参数化渲染的确定性摩尔机,并推出Tycho,这是一种编码智能体系统,可在交互过程中构建并使用特定游戏的模型。Tycho将可操作观测结果与中间动画、关卡完成及游戏结束帧区分开来,智能体可从该结构化历史中对自由形式的可执行假设进行建模、测试、规划、修复或绕过。在每项策略各进行一次匹配的公开集运行中,我们在匹配的推理预算下使用Claude Opus 4.8对全部25个公开游戏比较了四种编排策略:智能体请求委托给模型构建器的策略获得了观测到的最高平均相对人类行动效率(RHAE),为88.49;采用该选定策略后,GPT-5.6 Sol和Opus 5均达到100.00的RHAE并完成全部183个关卡,它们的游戏平衡首局人类复现中位排名分别为98.5和100.0,Opus 5的得分行动比官方人类基准的总和少61%;验证失败后的自动修复生成的模型能更准确地复现观测到的状态转移,但仅达到83.07的RHAE,状态转移匹配指模拟器复现观测动态的能力,而非其是否确定目标或改进下一步行动;强博弈还需决定何时构建、修复、使用或绕过模型,我们将这一联合问题称为主动抽象:从高成本交互中生成可测试模型,并判断何时获取或使用该模型值得付出成本。

英文摘要

ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal while maintaining action efficiency because every move counts. We formalize these environments as parameterized rendered deterministic Moore machines and introduce Tycho, a coding-agent system that constructs and uses game-specific models during interaction. Tycho separates actionable observations from intermediate animation, level-completion, and game-over frames. From this structured history, an agent can model, test, plan with, repair, or bypass a free-form executable hypothesis. In one matched public-set run per policy, we compare four orchestration policies on all 25 public games using Claude Opus 4.8 under matched inference budgets. Actor-requested delegation to a model builder obtains the highest observed mean Relative Human Action Efficiency (RHAE), 88.49. With this selected policy, GPT-5.6 Sol and Opus 5 both reach 100.00 RHAE and complete all 183 levels. Their game-balanced first-run human-replay midranks are 98.5 and 100.0. Opus 5 uses 61% fewer scored actions than the aggregate official human baselines. Automatic repair after verification failures produces models that reproduce observed transitions much more accurately, yet reaches only 83.07 RHAE. Transition match indicates whether a simulator reproduces observed dynamics, not whether it has identified the objective or improves the next action. Strong play also requires deciding when to construct, repair, use, or bypass a model. We call this joint problem active abstraction: generating a testable model from costly interaction and deciding when acquiring or using it is worth its cost.

Comments52 pages, 18 figures, 17 tables. Open-source implementation: https://github.com/NIMI-research/Tycho

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑