arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线策略蒸馏的收益与崩溃:一种强化学习视角

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue Zhang

arXiv 2610.03185首次发表:更新:

发表机构

Zhejiang University; Westlake University(浙江大学; 西湖大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从强化学习视角解释在线策略蒸馏(OPD)的性能提升与崩溃机制,发现教师隐式奖励放大学生行为,并通过屏蔽不健康响应和SFT初始化缓解崩溃。

AI 中文摘要

在线策略蒸馏(OPD)已成为语言模型后训练的重要方法。然而,尽管其带来了性能提升,OPD也可能崩溃为过度冗长和重复的生成,而对这些不同结果背后的机制仍知之甚少。我们通过强化学习的视角来解释这些结果:教师隐式地奖励学生的行为,即使这些行为教师自身也很少展现。从这个视角出发,我们的实验表明,OPD在不扩展学生能力的情况下提升了性能。当隐式奖励模型可靠时,OPD使正确响应的采样更容易。相反,当偏好与质量不一致时,就会发生奖励黑客行为:隐式奖励模型会放大过长、重复的学生轨迹,尽管教师自身很少生成此类文本。在这一诊断的指导下,我们发现,在训练期间屏蔽不健康的响应以及使用SFT初始化,都可以有效缓解崩溃。综合来看,这些发现表明,OPD放大了教师隐式反馈所青睐的学生行为,将焦点从教师生成的好坏转移到教师评估学生轨迹的可靠性上。我们的代码可在该https URL获取。

英文摘要

On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at https://github.com/HancCui/opd_hacking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑