发表机构
Shanghai Artificial Intelligence Laboratory; Harbin Institute of Technology; Fudan University; The Chinese University of Hong Kong(上海人工智能实验室; 哈尔滨工业大学; 复旦大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对强化学习中奖励监督问题,提出 H$^2$SD 混合事后自蒸馏框架。成功轨迹用教师概率调更新幅度,失败轨迹基于参考提示调整教师并最小化反向 KL 散度,实验表明该框架优于基线,能稳定优化且效率良好。
AI 中文摘要
具有可验证奖励的强化学习(RLVR)显著提高了大语言模型在数学推理和代码生成等任务上的推理能力。但多数 RLVR 方法给整个轨迹分配标量结果奖励,导致监督稀疏和令牌级信用分配受限。策略蒸馏(OPD)通过从更强的教师模型中蒸馏令牌级分布提供更密集监督,但需额外教师且通常假设词汇表共享。策略自蒸馏(OPSD)通过对特权信息调整同一模型构建教师策略来消除这种依赖,但直接匹配教师分布可能导致信息泄露和优化不稳定。RLSD 仅用教师信号调制更新幅度避免直接匹配,但采样推理失败时无法提供明确校正方向。为解决此权衡,我们引入 H$^2$SD,一种混合事后自蒸馏框架,根据轨迹正确性不同使用教师。对于成功轨迹,教师接收被确认为正确的学生响应及改写指令,其在原始响应令牌上的概率用于调制更新幅度而不改变奖励确定的方向。对于失败轨迹,我们基于包含关键推理步骤和已验证答案的参考提示调整教师,并最小化从学生到教师的反向 KL 散度。在多个具有挑战性的推理基准上的实验表明,H$^2$SD 始终优于代表性的 RLVR、OPSD 和 RLSD 基线,同时保持稳定优化和良好的生成效率。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teacher but typically assign it a fixed role: direct distribution matching may destabilize successful behavior, while magnitude-only modulation offers little corrective guidance after failure. We observe that successful and failed trajectories require different forms of hindsight supervision. A successful response already contains a valid student-generated reasoning path and can therefore serve as privileged context rather than being replaced by an external rationale. A failed response, however, requires corrective reference information. We introduce Hybrid Hindsight Self-Distillation ($\mathrm{H}^{2}\mathrm{SD}$), which jointly adapts teacher context and update strategy to trajectory correctness. For successful trajectories, we construct the teacher context from the verified response and a rephrasing instruction, and use the teacher only to re-evaluate the original response tokens. The resulting probabilities refine token credit assignment without changing the direction determined by the reward. For failed trajectories, a verified reference hint provides corrective guidance through reverse-KL distillation. Experiments on challenging reasoning benchmarks show that H$^2$SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.