发表机构
Salesforce AI Research; Salesforce Agentforce(Salesforce人工智能研究院; Salesforce Agentforce)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DarwinX是模型参数冻结时通过自然选择进化控制程序的方法,在四个基准中平均提升约17个百分点,控制程序可迁移,进化通用能力而非基准特定补丁。
AI 中文摘要
大语言模型(LLM)智能体的能力不仅取决于模型权重,还与其控制程序(harness)相关,控制程序包含提示词、工具、技能和控制流。现有自改进循环可编辑控制程序,但单谱系搜索具有路径依赖性,局部最优解常导致其他任务性能退化。本文提出DarwinX,该方法在模型参数冻结的情况下,将自进化过程视为对控制程序种群的选择:“保留并扩展”契约仅允许扩展覆盖范围且不发生退化的变体;档案库保存备选谱系以供重组;失败证据、教师证据和自生成证据共享同一编辑接口。适应度来自各基准自身的验证器,无需标准答案或人工挑选的优胜解。在四个逐步分离进化信号与测试的基准中,单轮循环平均提升约17个百分点:匹配基础模型时,Terminal-Bench 2.1提升7.7个百分点至83.2%,使用更强基础模型时达到验证前沿的84.7%;TerminalWorld的保留测试集达到68.3%,优于所有现成智能体;WebArena-Infinity真实任务的pass@1从43.5%提升至93.0%且符合审计要求;Terminal-Bench 2.1的控制程序可原封不动迁移至SWE-bench Verified。进化的是通用智能体能力,而非基准特定补丁,因此能在任务、验证器和基础模型变化时保持有效。冻结的模型无需是固定智能体:控制程序选择将评估计算转化为持久能力。
英文摘要
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.