arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合可验证奖励的在线蒸馏

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang

arXiv 2608.24696首次发表:更新:

发表机构

LeapLab, Tsinghua University; Beihang University; SMS, Peking University; NLPLab, Tsinghua University(清华大学LeapLab; 北京航空航天大学; 北京大学SMS; 清华大学NLPLab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出OPDVR方法,无缝结合OPD与RLVR且不增加超参数,经六个推理基准实验验证其性能优于标准OPD。

AI 中文摘要

带可验证奖励的强化学习(RLVR)与在线蒸馏(OPD)已成为大语言模型后训练的两种广泛采用的范式。然而,RLVR存在任务级反馈稀疏的问题,而OPD提供密集的token级指导却忽略轨迹正确性,其性能受限于教师模型。将二者结合是有前景的方向:OPD提供密集监督信号,RLVR提供任务级正确性。但现有整合常依赖加权组合或启发式切换,引入额外超参数与权衡。我们提出结合可验证奖励的在线蒸馏(OPDVR),一种简单有效的方法,无缝结合OPD与RLVR且不增加任何超参数。我们首先基于轨迹正确性重新表述采样token OPD的隐式奖励,然后应用ReLU门控机制确保正确轨迹获得非负奖励、错误轨迹获得非正奖励,从而使蒸馏信号与任务成功对齐,同时保留教师模型的分布指导。此外,我们的修改将采样token OPD转化为合适的RLVR方法,使其可轻松与任何策略梯度算法(如GRPO)结合。在六个推理基准上的实验表明,OPDVR始终优于标准OPD,代码可在该https链接获取。

英文摘要

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑