AI 中文总结
研究大语言模型智能体驾驭因素提升性能的问题,提出自我进化智能体驾驭框架,将提出与评估更改分开,通过门控存档避免过度拟合,在多领域测试中取得较好泛化效果,证明诊断与评估循环的有效性。
AI 中文摘要
大语言模型智能体在实际任务中的表现,不仅取决于其模型本身,还受围绕模型的驾驭因素影响,如提示、注入知识、运行时控制和配置等。在部署中,驾驭因素往往是提升性能的唯一途径。然而,自动改进驾驭因素面临挑战,因为自我生成的反馈有噪声,表面的提升可能是测量假象或过度拟合。本文提出了一个自我进化智能体驾驭框架,将提出更改与评估更改分开。语言模型诊断故障并提出补丁,而所有采样测量和显著性测试由确定性代码完成,确保改进的可信度。补丁填充基于编辑所解决的病理而非所修复任务的门控分类质量多样性存档,避免过度拟合。在七个领域对冻结的开放权重模型进行测试,训练选择并在密封测试中评分,所获提升为9到15.5个百分点,且保留了86%到147%的训练增益,证明了其泛化能力而非过度拟合。获胜补丁追踪模型的主要病理,而非其大小或家族。转移的是诊断与评估循环,而非任何特定的驾驭方式。
英文摘要
Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is governed by the surrounding agent harness, including prompts, tools, control loops, etc. Automatically evolving this harness offers a promising pathway to agent improvement, yet existing approaches typically rely on greedy candidate selection and noisy self-generated feedback, rendering their gains susceptible to search collapse, task-specific overfitting, and poor verifiability. To tackle these challenges, we introduce HarnessBank, a trustworthy agent-harness self-evolution framework that pairs a task agent with a separate evolver agent for iterative failure diagnosis, harness generation, and evolution verification. HarnessBank maintains a Harness Gene Bank composed of high-performing harnesses of different semantic coordinates. Those harnesses are reinvented, recombined, screened, and selected during the self-evolution procedure. Moreover, we propose a Gated Harness Screening mechanism to efficiently filter high-quality harnesses and reduce the cost of evaluating numerous offspring harnesses. Across seven agent benchmarks, HarnessBank produces consistent performance improvements from 5.1% to 15.4%. Cross-model experiments further verify that the improvements come from the model-specific self-evolving process, instead of a universally optimal harness. Our code will be publicly available upon acceptance.
Comments9 pages, 4 figures, 3 tables