arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28145cs.LG

RL始于RL之前:关于策略蒸馏以改进强化学习

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng… 展开作者

Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨同策略蒸馏作为强化学习准备阶段的价值,发现其能提升最终性能,且最佳蒸馏目标取决于轨迹来源和后续训练。

中文摘要 AI 辅助

强化学习(RL)能提升推理能力,但其性能取决于训练起始时的策略。我们将同策略蒸馏(OPD)作为RL的准备阶段进行研究,并探讨其益处是否仅限于提升蒸馏模型的初始准确率。在共享的RL设置下,使用OPD初始化的学生模型在最终性能上优于直接进行RL或先监督微调再RL的模型。即使OPD在准确率上几乎没有即时提升,这一优势也可能出现。RL前的Pass@k并不能完全解释这一益处:相似甚至更高的值并不必然带来RL后更好的性能。行为分析表明,与教师分布的对齐(超越top-1一致性)可能是潜在的解释。这种对齐可能有利于更高质量的推理路径,同时保留RL可利用结果反馈进一步优化的备选方案。我们进一步考察了轨迹来源和散度目标对后续RL蒸馏价值的影响。标准反向KL OPD在RL前表现更好,但前向KL OPD在RL后超越它;使用教师生成的蒸馏轨迹时,反向KL在两个阶段均保持领先。这些发现表明,偏好的蒸馏目标取决于轨迹来源和后续训练。我们的结果支持将OPD评估为RL的准备阶段,并根据后续训练后的性能来选择蒸馏方案。

英文摘要

Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward. Student rollouts outperform teacher rollouts under both objectives before and after RL. These findings highlight the importance of both objective choice and the states receiving supervision for subsequent RL. Our results support OPD as preparation for RL and favor forward KL when OPD is followed by RL in our comparison.

发表机构

  • Fudan University(复旦大学)
  • SII
  • Peking University(北京大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑