发表机构
Indian Institute of Science; Snap Inc.(印度科学理工学院; Snap公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DiffGate提出结果门控目标,结合GRPO与选择性有界教师指导,仅对失败轨迹按难度施加教师监督,在Qwen3模型上提升代码和数学的pass@8,改善解决方案覆盖。
AI 中文摘要
在线策略蒸馏(OPD)已成为大语言模型后训练中广泛使用的范式,通过让学生在自身生成的轨迹上接受监督,减少了传统蒸馏中训练与测试不匹配的问题。然而,现有的OPD目标在很大程度上仍是令牌局部且不考虑结果的,它们在每个前缀上优化师生一致性,尽管推理质量是在轨迹层面决定的。带可验证奖励的强化学习(RLVR),特别是组相对策略优化(GRPO),提供了互补的结果级监督,但存在奖励稀疏和信用分配粗糙的问题。我们表明,OPD和RLVR表现出互补的盲点:教师信号提供密集的局部指导,但与轨迹正确性弱对齐,而组相对奖励捕获任务成功,但提供粗粒度的令牌级信用,并在全失败组中消失。我们引入了DiffGate,一种结果门控目标,将GRPO与选择性、有界的教师指导相结合。教师监督仅应用于失败的轨迹,按组难度缩放,并平滑有界以防止极端的师生差异主导优化。因此,验证器决定哪些轨迹接收教师指导,而教师在这些轨迹内提供密集的令牌级更新方向。在Qwen3-0.6B和Qwen3-1.7B学生模型上,DiffGate在代码avg@8上分别比匹配的GRPO提高了+1.7和+1.8个百分点,pass@8提高了+1.6和+5.7个百分点。在数学上,avg@8保持在GRPO的0.5个百分点以内,而pass@8提高了+1.1和+3.9个百分点。总体而言,DiffGate在所有四个模型-领域设置中提高了pass@8,展示了在我们的评估协议下改进的解决方案覆盖范围。
英文摘要
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines \emph{which trajectories} receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by $+1.7$ and $+1.8$ points and pass@8 by $+1.6$ and $+5.7$ points, respectively. On mathematics, avg@8 remains within $0.5$ points of GRPO while pass@8 improves by $+1.1$ and $+3.9$ points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.
CommentsPreprint Under Review