当教师指导产生误导:奖励对齐的在线策略蒸馏
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
- State Key Laboratory of Novel Software Technology, Nanjing University(南京大学现代软件技术国家重点实验室)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对在线策略蒸馏中教师指导与结果奖励不一致的问题,提出RA-OPD方法,通过过滤不一致轨迹提升学生模型性能,在7个数学基准和3个代码基准上优于标准OPD及其他变体。
AI中文摘要:
在线策略蒸馏(OPD)近期成为大语言模型(LLMs)流行的后训练范式,为将教师模型的知识与能力迁移至学生模型提供了高效途径。然而,教师对学生生成前缀的指导并非总是可靠:训练本应优化模型生成更可能正确的响应,或等价于获得更高结果奖励,但OPD过程中,教师模型可能提供的指导会阻碍学生向正确轨迹移动,或使其向错误轨迹移动,这与结果奖励不一致。此类不一致的指导不可靠,会误导优化过程并最终降低模型性能。为缓解教师指导的不一致问题,我们提出奖励对齐的在线策略蒸馏(RA-OPD),其核心思路是仅保留那些能让学生向正确轨迹移动或阻止其向错误轨迹移动的轨迹。具体而言,对于每个采样轨迹,RA-OPD会检查其轨迹级蒸馏回报是否与结果奖励一致,随后过滤掉不一致的轨迹。RA-OPD选择更可靠的轨迹来提升学生模型性能,且无需额外计算成本。我们使用Qwen3系列和DeepSeek-R1系列模型在数学与代码基准上对RA-OPD进行评估,在7个数学基准和3个代码基准上,RA-OPD的性能显著优于标准OPD及其他测试的OPD变体。
英文摘要:
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidance on student-generated prefixes is not always reliable. Training should optimize the model to generate responses that are more likely to be correct, or equivalently, to get higher outcome rewards. But during OPD, the teacher model may provide guidance that discourages the student from moving toward correct trajectories or moves the student toward incorrect ones, which is misaligned with outcome reward. Such misaligned guidance is unreliable, as it would mislead the optimization process and ultimately degrade model performance. To mitigate misaligned teacher guidance, we propose Reward-Aligned On-Policy Distillation (RA-OPD). The key insight is to keep only trajectories whose induced updates move the student toward correct trajectories or discourage the student from moving toward incorrect ones. Specifically, for each sampled trajectory, RA-OPD checks whether its trajectory-level distillation return is consistent with its outcome reward and then filters out the misaligned trajectories. RA-OPD selects more reliable trajectories to improve student model performance without requiring additional computational cost. We evaluate RA-OPD on math and code benchmarks using models from the Qwen3 family and the DeepSeek-R1 family. Across seven math benchmarks and three code benchmarks, RA-OPD significantly outperforms standard OPD and other tested OPD variants.