arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于强化学习的智能体技能渐进式生成

Progressive Agent Skill Generation via Reinforcement Learning

Junhao Shen, Zhanqiu Zhang, Yiwen Guo, Hong Cheng

arXiv 2608.01678首次发表:更新:

发表机构

The Chinese University of Hong Kong; LIGHTSPEED(香港中文大学; 光速公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有技能生成方法的缺陷,提出基于强化学习的Skill-α方法,将技能生成为序列编辑过程并引入回退奖励,在CL-Bench和tau2-bench上均显著提升了下游任务成功率。

AI 中文摘要

现有的技能生成方法大多依赖启发式规则或流水线式整合,必须针对不同的证据源进行专门设计。相比之下,基于学习的方法提供了一种更统一的方式来对异构源的技能生成进行建模。然而,基于学习的技能生成仍然具有挑战性,因为技能缺乏基于相关性或正确性的自然监督信号;它们的价值很大程度上只能通过是否能提升智能体在下游任务上的表现来确定。为应对这一挑战,我们提出了Skill-α,一种用于渐进式生成高质量智能体技能的强化学习方法。具体而言,我们将技能生成为序列编辑过程,将技能构建分解为可单独评估的编辑操作,并引入了一种新颖的回退奖励机制,通过比较原始技能与编辑后的技能在锚定查询下的下游执行情况来评估每个编辑操作。大量实验表明,在文档到技能和经验到技能两种设置下,Skill-α生成的技能比基于启发式或流水线的方法更有效。在主GPT-4o工作智能体下,Skill-α在CL-Bench上的平均下游成功率比最强的技能生成基准提升了3.3个百分点,在tau2-bench上提升了6.7个百分点。进一步的 ablation 实验验证了回退奖励和渐进式生成的重要性。

英文摘要

Recent large language model agents often use external skills as modular procedural units that condition inference and improve complex task solving. Thus, automatically generating high-quality skills from documents or experience has become an important problem. Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches offer a more unified way to model skill generation across heterogeneous sources. However, learning-based skill generation remains challenging because skills lack a natural supervision signal based on relevance or correctness; their value can largely be determined only by whether they improve the behavior of the agent on downstream tasks. To address this challenge, we proposeSkill-$α$, a reinforcement learning method that learns a unified policy for progressive skill generation. Specifically, we construct each skill by repeatedly applying the learned policy to successive source evidence and introduce a novel rollback reward that evaluates each edit by comparing downstream execution under the original and edited skills on an anchored query. Extensive experiments show thatSkill-$α$ generates more effective skills than methods based on heuristics or pipelines in both document-to-skill and experience-to-skill settings. Under the main GPT-4o worker,Skill-$α$ improves average downstream success rates over the strongest skill-generation baseline by 3.1 points on CL-Bench and 6.7 points on tau2-bench. Further ablations and analysis validate the importance of rollback reward and progressive generation.

CommentsCode is available at https://github.com/ejhshen/skill-alpha

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑