arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI研究代理的递归自我改进

Recursive self-improvement of AI research agents

Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang

arXiv 2609.26457首次发表:更新:

发表机构

Weco AI(Weco AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出AIDE^2系统,实现AI研究代理的递归自我改进,通过自主修改代码并基准测试,在8天内发现七项改进,且收益泛化到未见任务,降低奖励黑客行为。

AI 中文摘要

AI代理正开始自动化AI栈中的研究与开发,从提高训练效率到优化推理。一个自然的下一步是提高代理自身的研究效率。当AI研究代理自身的代码成为优化对象时,每次被接受的改写都会成为下一轮编辑的代理。我们将此循环称为递归自我改进。其重要性在于一个长期趋势,即研发累计支出的增加带来递减的回报。持续的自我改进提供了一种对抗这一趋势的方法。我们提出了AIDE^2,一个为前沿AI研究代理实现此循环的系统。它对自己的代码提出修改建议,在一套AI研发任务上对修改后的自身版本进行基准测试,并保留在隐藏评估中表现最佳的修改。在一次自主的8天运行中,AIDE^2发现了七项连续的改进,从新的搜索策略到压缩和管理代理不断增长的上下文的记忆机制。这些收益泛化到四个保留的基准,涵盖机器学习工程、启发式算法工程和基于物理的天气预报,其中最后一个在分布外,与选择任务不同。在所有四个基准上,最强发现的代理达到或超过了人类工程化的生产研究代理,后者在FML-Bench上排名最强之一。在另一个独立的保留任务族上,发现的代理还表现出减少的奖励黑客行为,这是循环从未明确优化的属性:在运行期间,该比率从55%降至32%,比人类工程化代理低7个百分点。这些结果共同表明,AI研究代理可以通过递归自我改进提高自身的研究效率,并且这些收益转移到循环从未遇到的任务和领域。

英文摘要

AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each accepted rewrite becomes the agent that the next round edits. We refer to this loop as recursive self-improvement. Its significance lies in a long-standing trend, in which increased cumulative spending on R&D yields diminishing returns. Sustained self-improvement offers a way to counter this trend. We present AIDE^2, a system that implements this loop for a frontier AI research agent. It proposes changes to its own code, benchmarks modified versions of itself on a suite of AI R&D tasks, and keeps the changes that perform best on hidden evaluations. In an autonomous 8-day run, AIDE^2 discovered seven successive improvements, ranging from a new search policy to memory mechanisms that compress and manage the agent's growing context. These gains generalize to four held-out benchmarks spanning machine learning engineering, heuristic algorithm engineering, and physics-based weather forecasting, the last of which is out of distribution from the selection tasks. On all four, the strongest discovered agent matches or exceeds a human-engineered production research agent that ranks among the strongest on FML-Bench. On a separate held-out task family, the discovered agents also exhibit reduced reward hacking, a property the loop never explicitly optimized for: the rate falls from 55% to 32% during the run, 7 percentage points below the human-engineered agent. Together, these results show that an AI research agent can improve its own research efficiency through recursive self-improvement, and that these gains transfer to tasks and domains the loop never encountered.

Comments28 pages, 10 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑