arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.09304cs.CLcs.LG

SG-OPD: 通过符号一致性门控和分阶段教师采样的符号门控在线蒸馏

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

  • Zhejiang University(浙江大学)
  • Hunan University(湖南大学)
  • Tianjin University(天津大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoran Xu, Hongyu Wang, Yifei Gao, Jiaze Li, Xiaofeng Zhang, Xiaosong Yuan

更新

AI总结:

针对在线蒸馏中轨迹级对齐和教师偏好均匀可靠性假设的失效问题,提出SG-OPD方法,通过符号一致性门控和分阶段教师采样改进蒸馏效果,在竞赛级数学推理任务上平均提升1.98和7.50。

AI中文摘要:

在线蒸馏(OPD)在自身轨迹上训练学生模型,并利用更强教师的密集逐token监督,通常优于离线蒸馏和标准强化学习。然而,我们发现其有效性隐含地依赖于两个在实践中经常失效的假设:学生与教师之间的轨迹级对齐,以及教师偏好的均匀token级可靠性。因此,我们提出符号门控在线蒸馏(SG-OPD),该方法使用二元验证器作为教师信任信号,在两个互补粒度上发挥作用:分阶段教师采样在冷启动时混合验证器认可的教师轨迹,而符号一致性门控在教师与验证器校正方向一致的token上外推蒸馏更新,在分歧时内插。在竞赛级数学推理基准上的实验表明,SG-OPD持续优于标准OPD,在每样本和每问题水平上平均提升分别为1.98和7.50。

英文摘要:

On-policy distillation (OPD) trains a student on its own trajectories with dense per-token supervision from a stronger teacher, and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its effectiveness implicitly relies on two assumptions that frequently break in practice: trajectory-level alignment between the student and the teacher, and uniform token-level reliability of the teacher's preferences. We therefore propose Sign-Gated On-Policy Distillation (SG-OPD), which uses a binary verifier as a trust signal for the teacher at two complementary granularities: phased teacher sampling mixes in verifier-endorsed teacher rollouts at cold-start, and a sign-consistency gate extrapolates the distillation update on tokens where the teacher agrees with the verifier-correct direction and interpolates it where it disagrees. Experiments on competition-level mathematical reasoning benchmarks show that SG-OPD consistently outperforms standard OPD, with average gains of 1.98 and 7.50 at the per-sample and per-question levels, respectively.

↑