arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RRSI:智能体框架的正则化递归自我改进

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhuang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, Chen-Yu Lee

arXiv 2609.24972首次发表:更新:

发表机构

Google Cloud AI Research; Stanford University; Washington University in St. Louis; UNC-Chapel Hill(谷歌云AI研究院; 斯坦福大学; 圣路易斯华盛顿大学; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RRSI通过正则化约束智能体框架的递归自我改进,防止过拟合,在分布外基准上提升性能并减少30%策略令牌。

AI 中文摘要

大型语言模型智能体的能力在很大程度上由其框架(即围绕冻结骨干模型的提示、控制流、工具、记忆和上下文管理)所放大。近期方法通过迭代地提出和选择智能体框架的组件级编辑来自动化这一过程,实际上在智能体系统层面建立了一种递归自我改进(RSI)的形式。然而,这种递归进化可能通过记忆训练任务而过拟合,表现出在分布内的大幅提升,但在分布外基准上这些提升会缩小甚至消失。我们引入了智能体框架的正则化递归自我改进(RRSI),通过约束进化候选的提出和选择,将正则化原则融入框架自我改进中。提出者使用时间退火预算进行操作,限制候选可以捆绑的编辑数量,并基于进化历史鼓励未探索的轨迹。选择者配备了一个批评者和一个剪枝器:批评者筛选特定于基准的提议,而剪枝器移除那些过小、过昂贵或不再有用的更改。这些约束共同倾向于可复用的智能体机制,而非特定于基准的机制甚至噪声。在涵盖编码、智能体工作空间和工程设计任务的八个基准上,RRSI在其进化的分割上获得了高达14.1分的提升,在五个分布外基准上获得了高达4.7分的提升,同时产生的框架比未正则化的进化少使用30%的策略令牌。代码可在https URL获取,项目页面可在https URL获取。

英文摘要

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 6.0 points on the split it evolves against and up to 4.3 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑