自进化框架:以智能体自身为优化器在多任务上的应用
Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
浏览论文内容
中文总结 AI 辅助
提出自进化框架方法,让同一冻结模型既当求解器又当提议者,直接编辑自身框架,通过两阶段进化在多个基准上提升性能,并超越Codex。
中文摘要 AI 辅助
框架(harness)是围绕语言模型智能体的代码,用于组织提示、调用工具、管理上下文并控制执行。随着模型能力的增强,近期研究开始让智能体改进自身的框架,这一研究方向被称为自进化框架。在现有的大多数方法中,一个独立的提议者(proposer)在人工设计的框架上运行并修改求解器(solver)的框架,且每个基准测试都会单独进化一个框架。现实世界的任务来自多个领域,因此框架的进化和评估都应覆盖多样化的任务。我们提出了一个接近递归自我改进的框架:同一个冻结的模型,在相同版本的框架上,首先作为求解器解决任务,然后作为提议者,读取完整的运行记录并直接编辑运行该框架的代码。每轮进化从五个不同领域的基准测试中抽取任务。为了衡量泛化能力,训练任务和保留任务被严格分离,我们还额外在五个从未用于进化的分布外基准测试上进行评估。我们将进化过程视为深度学习训练,包含两个阶段:多任务预训练和持续训练。从49行的种子框架开始,第一阶段结束时获得的框架在分布内基准测试上将平均得分提高了4.48分,在分布外基准测试上提高了12.64分,在分布内基准上超越了Codex,在分布外基准上与Codex持平。在第二阶段,在分布外基准之一Claw-Eval上继续进化,将该基准的得分从66.17提高到68.06,超过了Codex。我们还对进化过程中出现的机制进行了深入分析,包括输出截断、历史压缩和独立审查。
英文摘要
A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks. We propose a framework close to recursive self-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held-out tasks are strictly separated, and we additionally evaluate on five out-of-distribution benchmarks never used during evolution. We frame the evolution process as deep-learning training with two stages, multi-task pretraining and continual training. Starting from a 49-line seed harness, the harness obtained at the end of the first stage improves the average score by 4.48 points on the in-distribution benchmarks and by 12.64 points on the out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw-Eval, one of the out-of-distribution benchmarks, further raises the score on that benchmark from 66.17 to 68.06, exceeding Codex. We also provide an in-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review.
发表机构
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。