arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33426cs.CL

TeacherGRPO:通过教师对齐弥合推理蒸馏中的能力差距

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, Jundong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对推理蒸馏中教师模型与学生模型能力差距导致的性能下降问题,提出基于强化学习的教师对齐方法TeacherGRPO,通过课程选择性对齐和重要性自适应长度正则化实现有效对齐,显著提升多种基准上的蒸馏性能。

中文摘要 AI 辅助

从强大的教师模型向较小学生模型进行推理蒸馏面临“差距诅咒”:随着教师模型变得更加复杂,其复杂分布与学生模型所能近似的分布之间的分歧越来越大,从而导致性能下降。现有的缓解策略要么通过数据选择过滤掉具有挑战性的示例,要么引入较弱的中间辅助模型,这本质上会损害监督覆盖范围或质量。我们提出教师对齐(Teacher Alignment),该方法直接使教师模型适应学生模型的分布,而无需丢弃数据或降低推理质量。然而,通过标准知识蒸馏进行的朴素对齐会引发教师模型推理能力的灾难性崩溃。为解决这一问题,我们将教师对齐重新表述为强化学习,并引入TeacherGRPO,它基于组相对策略优化(Group Relative Policy Optimization),并有两项关键创新:(i)课程选择性对齐(Curriculum Selective Alignment)应用双重令牌级和分布级课程,将奖励集中在高信号推理差距上,同时过滤来自琐碎令牌和不确定尾部分布的噪声;(ii)重要性自适应长度正则化(Importance-Adaptive Length Regularization)有选择地惩罚冗长冗余,同时保留教学上关键的推理步骤。对齐后的教师模型随后通过标准流程将知识蒸馏给学生模型。大量实验表明,TeacherGRPO在多种推理基准和蒸馏方法上显著优于基线方法。我们的代码可在以下网址获取:https://this URL。

英文摘要

Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • Netflix(网飞)
  • University of Washington(华盛顿大学)
  • Florida State University(佛罗里达州立大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑