arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DarwinX:通过自然选择进化智能体控制程序

DarwinX: Evolving Agent Harnesses Through Natural Selection

Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen

arXiv 2608.07545首次发表:更新:

发表机构

Salesforce AI Research; Salesforce Agentforce(Salesforce人工智能研究院; Salesforce Agentforce)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DarwinX是模型参数冻结时通过自然选择进化控制程序的方法,在四个基准中平均提升约17个百分点,控制程序可迁移,进化通用能力而非基准特定补丁。

AI 中文摘要

大语言模型(LLM)智能体的能力不仅取决于模型权重,还与其控制程序(harness)相关,控制程序包含提示词、工具、技能和控制流。现有自改进循环可编辑控制程序,但单谱系搜索具有路径依赖性,局部最优解常导致其他任务性能退化。本文提出DarwinX,该方法在模型参数冻结的情况下,将自进化过程视为对控制程序种群的选择:“保留并扩展”契约仅允许扩展覆盖范围且不发生退化的变体;档案库保存备选谱系以供重组;失败证据、教师证据和自生成证据共享同一编辑接口。适应度来自各基准自身的验证器,无需标准答案或人工挑选的优胜解。在四个逐步分离进化信号与测试的基准中,单轮循环平均提升约17个百分点:匹配基础模型时,Terminal-Bench 2.1提升7.7个百分点至83.2%,使用更强基础模型时达到验证前沿的84.7%;TerminalWorld的保留测试集达到68.3%,优于所有现成智能体;WebArena-Infinity真实任务的pass@1从43.5%提升至93.0%且符合审计要求;Terminal-Bench 2.1的控制程序可原封不动迁移至SWE-bench Verified。进化的是通用智能体能力,而非基准特定补丁,因此能在任务、验证器和基础模型变化时保持有效。冻结的模型无需是固定智能体:控制程序选择将评估计算转化为持久能力。

英文摘要

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑