arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillForge:通过动态技能生命周期协同进化技能与智能体

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Yuyao Ge, Yiwei Wang, Yuchen He, Baolong Bi, Lingrui Mei, Jiayu Yao, Lizhe Chen, Shenghua Liu

arXiv 2610.09832首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; University of California, Merced; Tsinghua University(中国科学院计算技术研究所; 加州大学默塞德分校; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillForge提出一种适应度驱动的技能生命周期方法,通过预退役、选择性退役和LLM引导变异使技能库与策略协同进化,在多个基准上取得最高成功率,相对改进达7.8%,并发布SkillFurnace数据集支持技能管理研究。

AI 中文摘要

记忆增强的强化学习增强了LLM智能体解决复杂长时程任务的能力。技能是这种记忆的一种形式,它将指令与任务类型上的适用性条件配对。然而,随着策略的改进,不加区分地保留每一个技能会让过时或有害的条目累积并误导智能体。我们提出了SkillForge,一种智能体强化学习方法,通过一个由试用、活跃、稳定和退役状态组成的适应度驱动的技能生命周期来编译和进化技能库,使得技能和模型在整个训练过程中协同进化。一个预强化学习评估阶段首先使用基础模型自身的轨迹来预先退役低适应度的技能,产生一个过滤后的库,该库随后用于监督微调的初始化。强化学习从这个检查点开始接管,并且在每次迭代中,选择性退役、稳定化和LLM引导的变异继续与策略优化一起锻造技能库。在多个交互式智能体基准测试中,SkillForge取得了最高的总体成功率,与最强基线相比提供了高达7.8%的相对改进,同时在整个训练过程中保持技能库紧凑。我们引入了SkillFurnace,一个包含5000多条带注释记录的数据集,这些记录捆绑了退役过滤的SFT轨迹、带有适应度注释的进化技能库,以及带有手动注释的失败类别的退役事件,以支持关于技能质量和生命周期管理的研究。

英文摘要

Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑