发表机构
AMAP, Alibaba Group(阿里巴巴集团AMAP)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillForge是实现技能验证与优化的持续技能演化框架,通过明确技能使用优化RL决策,在ALFWorld等数据集上优于SkillRL,可训练更强的LLM智能体。
AI 中文摘要
大型语言模型(LLM)智能体通过强化学习(RL)训练以完成复杂决策任务,但多数经RL训练的智能体仍为回合式,无法跨回合积累可复用知识。近期基于技能的方法如SkillRL尝试从原始轨迹提取技能,却将技能库视为仅可追加的存储库,未验证存储技能是否仍有效。本文提出SkillForge,一种用于持续技能演化的框架,可通过环境交互实现技能验证与优化。通过在智能体交互中明确技能使用,RL可直接优化环境动作与技能调用决策。SkillForge还引入基于证据的技能验证与多路径技能归纳,使技能库在持续增长的同时保持质量。在ALFWorld、WebShop与AppWorld上的大量实验表明,SkillForge始终优于SkillRL,证明持续验证的技能对训练更强LLM智能体的有效性。
英文摘要
Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumulate reusable knowledge across episodes. Recent skill-based approaches, such as SkillRL, attempt to address this issue by extracting skills from raw trajectories, but treat the skill bank as an append-only repository without verifying whether stored skills remain effective. In this paper, we propose SkillForge, a framework for continuous skill evolution that enables skills to be verified and refined through environment interaction. By making skill usage explicit during agent interaction, RL can directly optimize both environment actions and skill invocation decisions. SkillForge further introduces evidence-based skill verification and multi-pathway skill induction, allowing the skill bank to continuously grow while maintaining its quality. Extensive experiments on ALFWorld, WebShop, and AppWorld show that SkillForge consistently outperforms SkillRL, demonstrating the effectiveness of continuously verified skills in training stronger LLM agents.