发表机构
Peking University; University of Oklahoma; Tencent; Imperial College London; University of Michigan; University of Edinburgh; Xunce Technology; GienTech Technology(北京大学; 俄克拉荷马大学; 腾讯; 伦敦帝国理工学院; 密歇根大学; 爱丁堡大学; 讯策科技; 中电金信)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自进化智能体编辑技能文档时朴素验收门控导致永久退化且易受优化器诅咒的问题,提出统计验收门控SAGE,通过逐条目配对比较和单侧配对检验过滤不可靠编辑,在20个设置中19个降低退化率并全部达到最高分数。
AI 中文摘要
基于大语言模型(LLM)的智能体通过编辑持久化技能文档(该文档编码其工作流程、工具使用规则和决策逻辑)来自我进化。这一循环包含两个步骤:一个优化器提出候选编辑,一个门控接受或拒绝该编辑。以往工作集中于优化器,而门控仍遵循朴素规则,即保留任何能提高整体验证分数的编辑。我们证明该规则在两方面失效。首先,它允许永久性退化,因为编辑可能提高平均值却破坏技能已解决的条目。其次,它易受优化器诅咒影响,因为有限且有噪声的验证集上观测到的最佳分数存在向上偏差。为解决上述两个局限,我们提出用于自进化智能体的统计验收门控(SAGE)。与以往工作相比,SAGE 有两个贡献。第一,SAGE 提出逐条目配对比较,在相同验证条目上评估当前技能和编辑后技能,这能暴露整体分数所隐藏的退化,并对其进行非对称惩罚。第二,SAGE 采用单侧配对检验,仅当编辑的胜出在统计上可靠地超过其失败时才提交该编辑,否则弃权(不执行)。SAGE 是标准门控的保守改进,在边界设置下恰好恢复基线。它仅提交基线编辑的子集,过滤掉那些收益不可靠或通过破坏已解决条目而获得的编辑。在五个基准和四个骨干大语言模型上,在等预算协议下,SAGE 在 20 个设置中的 19 个中降低了退化率,在剩余一个设置中与基线持平,例如在 LiveMath 上从 36.5% 降至 0%,在 OfficeQA 上使用 DeepSeek-V4 时从 42.8% 降至 0%。SAGE 还在全部 20 个设置中达到最高最终分数,将 LiveMath 从 34.15 提升至 48.78。
英文摘要
Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.
Comments10 pages, 2 figures, 2 tables