发表机构
Tianjin University; Meituan(天津大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对GRPO与OPD简单结合的性能缺陷,提出RSTG方法,通过选择性自适应蒸馏及补充SFT,使数学任务性能提升4.02%、代码任务提升3.05%。
AI 中文摘要
带可验证奖励的强化学习(RLVR)已成为大型语言模型(LLM)后训练的标准范式。虽然组相对策略优化(GRPO)被广泛采用,但它存在奖励信号稀疏的问题,当组内所有响应获得相同奖励时会完全丢失梯度。在线策略蒸馏(OPD)通过提供教师模型的密集令牌级监督,成为一种自然的解决方案。然而,将GRPO与OPD简单结合会导致性能下降,其根本原因有三个:并非所有样本都能从蒸馏中受益;过快拟合教师会削弱RL的探索能力;OPD的优势不对称,会抑制大多数令牌。为解决这些挑战,我们提出RSTG(通过自适应教师指导恢复学习信号),它仅在最关键的地方选择性且精确地应用蒸馏。在样本层面,OPD仅应用于负零方差提示,每个样本按教师的置信度评分加权;在令牌层面,蒸馏仅针对学生熵高或师生分歧大的令牌。我们还通过对教师模型生成的正确轨迹进行监督微调(SFT)来增强训练,在RL无法产生梯度的地方注入正梯度信号。实验表明,RSTG在数学任务上比简单的GRPO+OPD组合性能提升4.02%,在代码任务上提升3.05%。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals and loses gradients entirely when all responses within a group receive identical rewards. On-policy distillation (OPD) offers a natural remedy by providing dense, token-level supervision from a teacher model. However, naively combining GRPO with OPD leads to degraded performance, due to three underlying causes: not all samples benefit from distillation; fitting too quickly to the teacher undermines the exploratory capacity of RL; and OPD's advantages are asymmetric, suppressing most tokens. To address these challenges, we propose RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), which applies distillation selectively and precisely where it matters most. At the sample level, OPD is restricted to negative zero-variance prompts with each sample weighted by the teacher's confidence score. At the token level, distillation targets only tokens with high student entropy or large teacher-student divergence. We further augment training with SFT on correct trajectories generated by the teacher model, injecting positive gradient signals where RL yields none. Experiments demonstrate that RSTG substantially outperforms naive GRPO+OPD by +4.02% on math and +3.05% on code.