发表机构
Tsinghua University; Northwestern Polytechnical University(清华大学; 西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillCycle通过技能内化与规则修订的反馈循环协同进化智能体策略和技能库,在WebShop上以3B模型取得74.74%成功率,显著优于静态技能库。
AI 中文摘要
将外部技能内化会改变语言智能体的能力,并随之改变其剩余指导的价值:规则可能变得冗余、具有误导性或不足以应对新遇到的决策。这产生了一个耦合问题,即从技能中学习以及调整监督进一步学习的技能。我们引入了SkillCycle,这是一个通过技能内化与规则修订之间的反馈循环来协同进化智能体策略和技能库的框架。我们的核心贡献是赋予蒸馏反馈第二个角色:token级别的上下文差异有助于定位需要检查的规则,而交互结果则指导对其内容和适用性的编辑。SkillCycle在两个阶段之间交替进行:使用固定技能库和路由器的策略学习,以及使用冻结策略的规则修订。候选编辑在指导下一个学习周期之前,会经过规则级和整个技能库的环境比较。在WebShop上,使用3B模型的SkillCycle在无需推理时技能输入的情况下达到了74.74%的成功率和88.37的得分,相对于最先进(SOTA)模型分别实现了0.73%和3.96%的相对改进。在ALFWorld和WebShop上的Cycle 3消融实验中,SkillCycle的无技能成功率相对于静态技能库分别提高了10.18%和18.11%,相对于单次技能库更新分别提高了2.41%和2.50%。这些结果表明,随着智能体能力的变化而持续修订技能指导,有助于将外部技能转化为推理时无需技能输入的策略能力。我们将发布代码、配置、技能库和评估协议。
英文摘要
Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupled problem of learning from skills and adapting the skills that supervise further learning. We introduce SkillCycle, a framework for co-evolving agent policies and skill banks through a feedback loop between skill internalization and rule revision. Our central contribution is to give distillation feedback a second role: token-level contextual differences help locate rules for inspection, while interaction outcomes guide edits to their content and applicability. SkillCycle alternates between two phases: policy learning with a fixed skill bank and router, and rule revision with a frozen policy. Candidate edits undergo rule-level and whole-bank environment comparisons before they guide the next learning cycle. On WebShop, SkillCycle with a 3B model achieves a success rate of 74.74% and a score of 88.37 without inference-time skill inputs, representing relative improvements of 0.73% and 3.96% over the state-of-the-art (SOTA) model, respectively. In Cycle 3 ablations on ALFWorld and WebShop, SkillCycle's no-skill success rates improve by 10.18% and 18.11% relative to a static skill bank, and by 2.41% and 2.50% relative to a single bank update, respectively. These results show that continually revising skill guidance as the agent's capabilities change helps transform external skills into policy capabilities that require no skill inputs at inference. We will release code, configurations, skill banks, and evaluation protocols.