AI 中文总结
该研究发现LLM智能体自进化会出现技能污染导致性能下降的不可逆问题,提出VaG预提交门控机制,可提升性能并实现技能池跨模型迁移。
AI 中文摘要
自进化智能体通过从执行轨迹中提炼可复用技能来积累能力,但研究发现这一过程并非单调:当技能池规模超过临界值后,新增技能会降低性能而非提升。我们将这种能力污染相变形式化,并追溯其结构成因:一旦有缺陷技能进入决策上下文,就会成为后续技能提炼的参考材料,形成跨轮污染链。进一步研究表明,这种污染在结构上是不可逆的:事后移除源技能无法消除其后代已继承的错误推理,因此事后回滚仅能恢复小部分损失性能。这使得技能准入成为预提交的必要条件,而非事后补救,进而提出了“验证器作为守门人(Verifier-as-Gatekeeper, VaG)”机制:一种渐进式信任层级,包含三类异构评判器——结构有效性、行为无害性和语义一致性,分别对每个技能进行单独过滤,同时搭配边际增益子集选择,在技能进入运行时上下文前,于顶层消除组合式污染。在Terminal-Bench 2上,无条件积累性能先升至峰值后下降,随着技能池扩大,大部分增益被抵消,事后移除罪魁祸首技能仅能恢复小部分性能下降,这是不可逆性的经验特征。相比之下,VaG性能每轮均提升,在规模约小5倍的技能池下达到72%的pass@1,且其冻结技能池无需重新进化即可正向迁移至其他四个主干模型及第二个基准测试。消融实验证实三类评判器互补且不可相互替代,每类拦截的有害技能类别基本不重叠。
英文摘要
Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it. We formalize this capability-contamination phase transition and trace it to a structural cause: once a defective skill enters the decision context, it becomes reference material for distilling later skills, forming cross-round contamination chains. We further show the contamination is structurally irreversible: removing a source skill after the fact cannot erase the flawed reasoning its descendants have already inherited, so post-hoc rollback recovers only a small fraction of the lost performance. This makes skill admission a pre-commit necessity rather than a post-hoc fix, and motivates Verifier-as-Gatekeeper (VaG): a progressive trust hierarchy whose three heterogeneous critics - structural validity, behavioral harmlessness, and semantic consistency - filter each skill individually, coupled with a marginal-gain subset selection that removes combinatorial contamination at the top tier before skills reach the runtime context. On Terminal-Bench 2, unconditional accumulation rises to a peak and then degrades, giving back most of its gains as the pool keeps growing, and post-hoc removal of the culprit skills recovers only a small part of the drop - the empirical signature of irreversibility. In contrast, VaG improves every round, reaching 72% pass@1 with a pool roughly 5x smaller, and its frozen skill pool transfers positively to four other backbones and a second benchmark without re-evolution. Ablations confirm the three critics are complementary and mutually non-substitutable, each intercepting a largely disjoint class of harmful skills.