arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PR-OPD:用于智能体强化学习的特权表示在线策略自蒸馏

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, Shigang Chen

arXiv 2609.36642首次发表:更新:

发表机构

University of Florida; University of Illinois at Chicago; Case Western Reserve University; Vanderbilt University; Weill Cornell Medicine(佛罗里达大学; 伊利诺伊大学芝加哥分校; 凯斯西储大学; 范德堡大学; 威尔康奈尔医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PR-OPD,利用技能引起的隐藏状态变化进行在线自蒸馏,在ALFWorld和WebShop上超越GRPO,最高提升14.0个百分点。

AI 中文摘要

语言模型智能体通常通过每个回合一个奖励的强化学习进行训练,而特权自蒸馏通过让同一策略在拥有技能时,通过令牌概率将其技能传授给无技能的自身,从而丰富了这一训练过程。然而,我们发现了两个质疑这一渠道的现象。不可见优势:上下文中的技能将WebShop成功率从42.2%提升至56.2%,却仅改变了不到四分之一的采样令牌的概率。大量对齐:技能改变了超过80%的响应令牌的隐藏状态,且线性探针可以追溯到特定技能。为利用这一点,我们提出了特权表示在线策略自蒸馏(PR-OPD)。在GRPO热启动后,策略为每个轨迹编写事后技能,以该技能作为停止梯度教师重新阅读自身响应,并在每一层将投影后的隐藏状态与教师的隐藏状态对齐,同时优化奖励目标,无需外部技能库、独立教师或推理开销。在ALFWorld和WebShop上使用两种骨干网络,PR-OPD在所有设置中均取得了最佳总体结果,相比GRPO在ALFWorld成功率上最高提升4.7个百分点,在WebShop准确率上最高提升14.0个百分点。代码可在该https URL获取。

英文摘要

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hindsight skill for each trajectory, re-reads its own responses with that skill as a stop-gradient teacher, and aligns its projected hidden states to the teacher's at every layer alongside the reward objective, with no external skill library, separate teacher, or inference overhead. On ALFWorld and WebShop with two backbones, PR-OPD achieves the best overall results in every setting, improving over GRPO by up to 4.7 points in ALFWorld success and 14.0 points in WebShop accuracy. Code is available at https://github.com/balibata/PR-OPD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑