arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15439cs.AI

解决ARC-AGI-3编码智能体是否需要可执行世界模型、简化和验证?

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

Sergey Rodionov

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨解决ARC-AGI-3时编码智能体是否需可执行世界模型、简化和验证。通过四个基于Codex的嵌套智能体评估,发现各变体随模型和推理强度改进,组件影响因设置而异,完整验证处理最佳,文本变体在部分设置中表现出色。

中文摘要 AI 辅助

我们之前的ARC-AGI-3智能体集成了可执行世界建模、计划简化和精确重放验证,不清楚哪种理念导致其性能提升。本文用四个基于Codex的嵌套智能体解决归因问题:文本基线;无重放验证的灵活接口可执行世界模型;有计划简化的相同可执行模型;保留简化并要求精确再现记录观测的固定接口验证处理。主要研究在公共ARC-AGI-3游戏中用gpt-5.4和gpt-5.5对四个智能体进行高和极高推理强度评估。探索性后续研究用gpt-5.6-sol在极高和最大推理强度下评估文本和验证变体。最显著结果是每个智能体变体都随更强模型和更大推理强度而改进。在每个模型强度设置中,变体间差异小于预期,而单个组件的影响因设置而异。要求持久可执行交付物并非普遍有益:文本变体在gpt-5.5的两种设置中都优于灵活接口可执行变体。简化在四个模型强度设置中的三个中提高了性能,最弱设置除外。完整验证处理在所有四个设置中排名第一,尽管使用了更多资源。在gpt-5.6-sol后续研究中,验证变体在两种推理强度下都完全解决了每个公共游戏,实现了约99%的RHAE,且使用的总动作少于人类基线的一半。由于模型晚于这些游戏,且未测试留出的性能,此结果应仅解释为公共集的饱和。

英文摘要

Our previous ARC-AGI-3 agent bundled executable world modeling, prompted simplification, and exact replay verification, leaving their individual contributions unclear. An executable world model is a persistent, agent-authored environment hypothesis embodied in runnable code. We compare four Codex-based variants: textual; flexible-interface executable; executable with simplification prompts; and a fixed-interface variant with simplification and exact replay verification against recorded observations. The main study evaluates them with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on 25 public games; exploratory follow-ups compare textual and verification with gpt-5.6-sol. In the main study, every variant scores higher as model capability and reasoning effort increase. These gains often exceed variant differences, which are smaller than anticipated and vary across settings. Requiring an executable deliverable is not universally beneficial: textual outperforms flexible-interface executable in both gpt-5.5 conditions. The simplification variant scores higher than its executable-only counterpart in three of four settings; the weakest is the exception. The complete verification treatment ranks first throughout, sometimes narrowly, but uses substantially more resources. With gpt-5.6-sol, the verification variant completes every public level at xhigh and max with about 99% human-relative action efficiency while using fewer than half the human baseline's total actions. At max, however, the textual variant completes every level with 41% fewer actions than the human baseline. Thus, at max, the three imposed mechanisms are not required for action-efficient public-set completion; verification nevertheless scores higher and succeeds at lower effort. Because gpt-5.6-sol postdates the games and held-out performance is untested, results indicate public-set saturation only.

补充信息

↑