重新思考自进化:缓解技能过拟合的约束探索-利用过程
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
浏览论文内容
中文总结 AI 辅助
针对LLM智能体技能过拟合问题,提出SkillBoost三阶段框架,在23种模型-基准配置上达SOTA性能,优化后技能可跨智能体复用。
中文摘要 AI 辅助
让大语言模型(LLM)智能体从过往交互中积累并复用经验,是现实应用中的核心挑战。一种有前景的解决方案是将技能视为可训练状态,按神经网络训练中优化模型参数的方式优化这些技能。然而,数据驱动的技能优化易过拟合于真实环境中收集的有限轨迹:过度利用这些轨迹会过拟合当前批次,而无约束的探索则会导致已解决案例的性能退化。这种权衡关系催生了对技能自进化的约束搜索视角,由探索-利用权衡主导。我们提出SkillBoost,这是一个缓解上述两种风险的三阶段框架:结构化利用将观测到的故障定位到可编辑的技能组件;先验引导的探索利用LLM中的先验知识生成多样的修复候选;验证接受机制仅在候选能在退化边界内提升性能时才采用该候选。在23种模型-基准配置上的实验表明,SkillBoost在缓解过拟合的同时达到了SOTA性能,优于人工构建和LLM生成的技能;迁移实验进一步显示,优化后的技能可被其他智能体在相似任务上复用。
英文摘要
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases. This tension motivates a constrained search view of skill self-evolution, governed by an exploration--exploitation trade-off. We propose SkillBoost, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound. Experiments across 23 model--benchmark configurations show that SkillBoost achieves state-of-the-art performance while mitigating overfitting, outperforming both human-crafted and LLM-generated skills. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks.
发表机构
- Zhejiang University(浙江大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。