arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从模仿到奖励发现:智能体强化学习的同策略预热

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu

arXiv 2609.39436首次发表:更新:

发表机构

Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出同策略预热(OPW)方法,通过教师监督学生自身交互轨迹进行蒸馏,为智能体RLVR提供有效初始化,理论证明其加速奖励发现并提升最终性能。

AI 中文摘要

带可验证奖励的强化学习(RLVR)为训练语言模型智能体提供了一种可扩展的方法,然而稀疏的结果奖励可能在训练早期为策略改进提供很少的信号。我们识别出一种同策略加速现象:在我们的主要比较中,以同策略蒸馏初始化的RLVR在训练早期就达到较高性能,并且在后续RLVR期间取得更高的平均性能以及比替代基线更高的最终性能。受此观察启发,我们研究了同策略预热(OPW),这是一个教师引导的阶段,在该阶段中,学生在其自身的交互轨迹上接受教师监督训练,然后再过渡到RLVR。与在固定的教师生成轨迹上进行模仿不同,OPW针对的是由学生自身决策引发的状态,包括不完美的动作和恢复情境。我们通过将同策略反向KL蒸馏与轨迹级分布匹配联系起来,提供了一种理论解释。在教师能力较强且总体蒸馏损失足够小的情况下,这种联系为初始验证器成功率提供了一个下界,并相应地给出了奖励发现复杂性的界。对于群体相对RLVR,我们进一步刻画了何时增加的成功概率会产生更具奖励信息量的群体。总之,我们的研究结果支持同策略蒸馏作为智能体RLVR的有效预热方法,并确定初始奖励发现是可能促成所观察到的加速现象的一种机制。

英文摘要

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑