AI 中文总结
本研究提出迁移分层在线策略蒸馏(TS-OPD),通过按学生与教师联合成功筛选问题并路由至门控前向/反向KL,在数学推理中实现强化学习改进教师对学生模型的结构化迁移,提升宏平均正确率。
AI 中文摘要
强化学习可以显著改进推理教师模型,但尚不清楚当该教师监督一个较小的在线策略学生模型时,这些改进中有哪些能够保留。我们在数学推理中通过比较GRPO前后的教师谱系、多种学生规模、直接GRPO以及多种在线策略蒸馏目标来研究这一问题。核心发现是迁移是结构化的而非标量性的:仅凭教师强度并不能使密集蒸馏具有竞争力,而强化学习改进的教师能带来有用但依赖于指标的学生收益。这促使我们提出迁移分层在线策略蒸馏(TS-OPD),该方法通过学生和教师联合采样的成功情况来筛选训练问题,将获取问题路由到门控前向KL,将巩固问题路由到门控反向KL,并添加熵制动以保护采样覆盖。在主要比较中,对于使用GRPO改进教师的宏平均正确率,TS-OPD是最强的学生目标,而pass@K则仍然较为混杂。消融实验表明,收益来自路由和令牌门控,而非跳过问题。这些结果支持一种迁移感知的在线策略蒸馏观点:当监督方向和令牌预算与学生观察到的能力相匹配时,更强的教师会有所帮助,而不仅仅是因为教师端模型更强。
英文摘要
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.