arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30861cs.AI

SkillEvoReg:针对过拟合的智能体技能进化正则化

SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

针对语言模型智能体技能进化中的过拟合问题,提出通用正则化框架SkillEvoReg,结合技能丢弃、复杂度正则化与因果反例验证,在多个基准上控制技能增长并提升下游能力。

中文摘要 AI 辅助

语言模型智能体越来越倾向于通过将执行经验转化为可重用的外部技能来提升自身能力。然而,反复的技能更新本身构成了一个学习过程:局部有用的编辑可能累积成冗余或任务特定的指令,而新的更新可能破坏先前有效的行为。我们将此问题研究为技能进化过拟合,并引入SkillEvoReg,这是一个受神经网络训练中抗过拟合技术启发的通用技能进化正则化框架。SkillEvoReg结合了训练时的技能丢弃(扰动更新生成)、复杂度感知的局部正则化(控制不必要的结构增长),以及因果反例验证(CCV,提供针对候选特定回归的定向行为验证)。我们在异构技能进化系统中实例化该框架,同时保留各系统原生的技能进化器和任务评估器。在SkillOpt、SkillEvolBench和ContinualSkillBench上,SkillEvoReg持续控制技能状态增长,同时保持有竞争力的下游能力,改善了几项迁移和后期进化结果,并识别出仅靠结构指标无法揭示的更新级回归。这些结果表明,显式正则化是对日益强大的技能更新器的有益补充。

英文摘要

Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neural-network training. SkillEvoReg combines training-time skill dropout, which perturbs update generation, and complexity-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation (CCV), which provides targeted behavioral validation of candidate-specific regressions. We instantiate the framework across heterogeneous skill-evolution systems while retaining each system's native skill evolver and task evaluator. Across SkillOpt, SkillEvolBench, and ContinualSkillBench, SkillEvoReg consistently controls skill-state growth while preserving competitive downstream capability, improves several transfer and later-stage evolution outcomes, and identifies update-level regressions that structural metrics alone cannot reveal. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters.

↑