发表机构
AWS(亚马逊云科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过追踪编码智能体玩ARC-AGI-3时的文件痕迹,提出测量协议分析其持续学习机制,发现其选择性遗忘等特性,为持续学习提供经验。
AI 中文摘要
我们研究编码智能体如何在一系列抽象推理任务中学习。该智能体运行于固定框架内的冻结基础模型上,通过编写和运行Python及shell脚本执行动作;除了其编写的产物外,它在各轮之间不保留任何状态,因此其形成、携带、修正或放弃的每一个思考都会留下痕迹,其中思考指的是任何写入文件的信念、规则或计划。我们让该智能体玩ARC-AGI-3,这是一组不提供任何说明的交互式推理游戏,每个游戏由一系列关卡组成,清除某一关卡的策略可能在下一关卡失效,因此每个新关卡实际上是一项新任务。智能体将其所学记录为Python脚本和文本笔记,而框架则保留所有动作和观察的完整日志。我们的贡献是一种测量协议,该协议通过这些文件追踪每个思考,从其形成的任务到被修正或放弃的任务,该协议被应用于来自两个模型家族的三个主干模型的七次评估运行。为一项任务编写的脚本几乎不会在后续任务中再次被调用(630次引用中仅有33次跨任务边界),因为大多数脚本嵌入了当前关卡的状态;相反,智能体将其知识重写为新脚本,保留通用规则并丢弃特定于关卡的细节,且会放弃其在边界前编写的74%的脚本。仅由模型读取的笔记从未被修正:智能体仅追加内容而不删除早期声明,积累的矛盾会根据日志解决。由于日志保留了所有内容,智能体是选择性遗忘而非灾难性遗忘。最代价高昂的错误是硬编码值被带入不再适用的任务。这些发现来自智能体编写的文件,无需访问模型,构成了对编码智能体如何持续学习的白盒分析。
英文摘要
We study how a coding agent learns across a sequence of abstract reasoning tasks. The agent runs on a frozen foundation model inside a fixed harness and acts by writing and running Python and shell scripts. It retains no state across turns other than its written artifacts, so every thought it forms, carries, corrects or abandons leaves a trace, where a thought is any belief, rule or plan committed to a file. We let the agent play ARC-AGI-3, a set of interactive reasoning games that provide no instructions. Each game is a sequence of levels, and a strategy that clears one level can fail on the next, so every new level is in effect a new task. The agent records what it learns as Python scripts and text notes, while the harness keeps a complete log of every action and observation. Our contribution is a measurement protocol that traces each thought through these files, from the task where it forms to the task where it is corrected or abandoned, applied to seven evaluation runs with three backbones from two model families. Scripts written for one task are almost never called again in a later task (33 of 630 references cross a task boundary), because most scripts embed the state of the current level. Instead, the agent rewrites its knowledge into new scripts, keeping the general rules and dropping the level-specific details, and abandons 74% of the scripts it wrote before a boundary. The notes, which only the model reads, are never revised: the agent appends without removing earlier claims, and the contradictions that accumulate are settled against the log. Because the log preserves everything, the agent forgets selectively, not catastrophically. The most costly error is a hard-coded value carried into a task where it no longer holds. These findings come from the files the agent wrote, without access to the model, and constitute a white-box analysis of how a coding agent continually learns.
CommentsAccepted at the NeurIPS 2026 Workshop on Continual Learning in the Era of Foundation Models and Embodied Agents (CL4FMAgents)