arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoSkill:用于层级技能演化的推理智能体与元技能智能体的联合强化学习

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

Jinyuan Feng, Dongmin Li, Yiqun Chen, Yang Gao, Xing Chen, Huimu Wang, Zhiqiang Pu

arXiv 2609.04865首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; Renmin University of China; ByteDance(中国科学院自动化研究所; 中国人民大学; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoSkill是一个统一多智能体RL框架,将静态元技能工作流转为可学习的元技能智能体,与推理智能体在层级技能库上联合训练,在ALFWorld和WebShop上的成功率分别达98.4%、90.6%,优于现有基线。

AI 中文摘要

技能库通过让大语言模型(LLM)智能体重用程序性知识,提升了智能体强化学习(RL)的样本效率。但现有范式存在结构缺陷:要么将技能演化与策略优化解耦,要么将元技能实例化为固定工作流,两者均把技能视为待管理的被动对象,限制了技能的灵活演化及其与推理智能体的协同适应。为解决这些局限,本文提出CoSkill,这是一个统一的多智能体RL框架,将静态元技能工作流重构为可学习的元技能智能体,并在层级技能库上与推理智能体联合训练。通过将推理智能体和元技能智能体建模为共享单一主干的协作团队,CoSkill实现了端到端协同适应:推理智能体根据检索到的任务技能及从其子集中选的步骤技能来确定动作,而其任务表现则指导元技能智能体优化这些步骤技能。在ALFWorld和WebShop上的实验表明,CoSkill的表现显著优于现有基于技能的基线和RL基线,分别达到98.4%和90.6%的成功率(提升3.5和6.2个百分点)。如图1所示,CoSkill在早期样本效率、渐近性能和挂钟效率上均表现更优。我们的代码可在该https URL获取。

英文摘要

Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple skill evolution from policy optimization or instantiate meta-skills as fixed workflows. Both treat skills as passive objects to be managed, limiting the flexible evolution of skills and their co-adaptation with the reasoning agent. To address the limitations, we propose CoSkill, a unified multi-agent RL framework that recasts the static meta-skill workflow as a learnable Meta-Skill Agent and jointly trains it with a Reasoning Agent over a hierarchical skill library. By modeling the Reasoning and Meta-Skill Agents as a cooperative team sharing a single backbone, CoSkill enables end-to-end co-adaptation: the Reasoning Agent conditions its actions on a retrieved task skill and step skills selected from its child set, while its task performance guides the Meta-Skill Agent in refining those step skills. Experiments on ALFWorld and WebShop show that CoSkill substantially outperforms prior skill-based and RL baselines, achieving success rates of 98.4% and 90.6%, respectively (+3.5 and +6.2 pp). As shown in Figure 1, CoSkill achieves superior early-stage sample efficiency, asymptotic performance, and wall-clock efficiency. Our code is available at https://github.com/jinyuan-cookie/CoSkill.

CommentsWithdrawn pending internal content review and approval by the authors' institution. An updated version will be resubmitted once the approval process is completed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑