arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RetireOPD:面向智能体强化学习的自退役在策略蒸馏

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Qianglong Chen, Yongliang Shen

arXiv 2609.20784首次发表:更新:

发表机构

Zhejiang University; Alibaba Group(浙江大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RetireOPD通过自适应退役机制,在智能体强化学习中利用技能条件化教师进行在策略蒸馏,显著提升ALFWorld和WebShop任务性能。

AI 中文摘要

使用强化学习(RL)训练的多轮智能体每个轨迹仅获得一个标量奖励,这促使自蒸馏(OPD)通过具有特权任务技能的自教师提供密集的令牌级监督,使无技能学生能够内化这些技能。然而,这一方法在智能体任务中受到两个发现的削弱:仅凭特权信息并不总能保证教师的可靠性,且教师监督的益处具有阶段依赖性。因此,我们提出RetireOPD(自退役在策略蒸馏),该方法首先使用环境奖励优化一个解耦的、技能条件化的教师,然后联合使用RL和OPD训练一个无技能学生。RetireOPD并未遵循预定义的蒸馏时间表,而是采用自适应退役:一旦学生与教师之间的差异停止缩小,且学生达到教师成功率的预定比例,学生便自行放弃教师,此后仅使用RL继续训练。在从1.5B到7B的Qwen2.5模型中,RetireOPD将ALFWorld的成功率相对于RL基线提高了14.1%至18.8%,将WebShop的准确率提高了11.8%至19.0%,并且在所有设置中均超越了其自身的技能条件化教师。

英文摘要

Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑