发表机构
AWS AI Labs; Georgia Institute of Technology(亚马逊云科技人工智能实验室; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FREEEVOLVE提出一种自动化智能体进化过程的方法,通过元进化学习可编辑的进化技能,在多个基准上平均提升13.6点,并展示跨环境迁移能力。
AI 中文摘要
智能体进化器(Agent evolvers)自动化了围绕语言模型智能体的提示、技能和工作流的设计,然而它们所遵循的优化过程仍然是手工设计的:一个固定的搜索循环决定候选如何评估、哪些被保留以及搜索何时停止。我们提出了FREEEVOLVE,它也自动化了这一过程。环境指定了目标、目标智能体、评估器、数据和资源限制;在这些限制内,进化器自己决定测试什么、收集多少证据、追求哪些候选以及何时停止。这些决策遵循一个可编辑的进化技能,我们通过元进化来改进它,即对每个候选技能在其产生的新目标智能体上进行评分。因此,优化过程变成了一种从经验中学习的能力,而不是预先设计的循环。在tau3-bench、ARC-AGI-2、ARC-AGI-3和Terminal-Bench 2.1上,FREEEVOLVE自行控制进化活动,但平均将主要保留指标提高了13.6个百分点,并达到或超过了手工设计的进化器。学习到的过程随着经验不断提高:元进化技能在新目标智能体上比种子技能增加了6.9个百分点,展示了跨环境的可迁移性。
英文摘要
Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this process as well. An environment specifies the goal, target agent, evaluator, data and resource limits; within these limits, the evolver itself decides what to test, how much evidence to collect, which candidates to pursue and when to stop. These decisions follow an editable evolution skill, which we improve through meta-evolution by scoring each candidate skill on the fresh target agent it produces. The optimization process thus becomes a capability learned from experience rather than a loop engineered in advance. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1, FREEEVOLVE controls the evolution campaign by itself, yet improves the primary held-out metric by 13.6 points on average and matches or exceeds hand-designed evolvers. The learned process keeps improving with experience: meta-evolved skills add 6.9 points over the seed skill on fresh target agents, demonstrating transferability across environments.