arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

传递接力棒:轨迹中继的在线策略蒸馏

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni, Hongxing Li, Yiwen Qiu, Weiming Lu, Yongliang Shen

arXiv 2607.26057首次发表:更新:

发表机构

Zhejiang University; Alibaba Group(浙江大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对在线策略蒸馏的前缀失败问题,利用师生延续不对称性提出Relay-OPD方法,通过教师在触发点接管生成轨迹段,让学生在其上优化,在数学推理基准测试中取得优异成绩,提升效果并减少训练轨迹长度。

AI 中文摘要

在线策略蒸馏(OPD)在学生自身轨迹中进行令牌级监督,但存在前缀失败问题:一旦学生陷入错误推理方向,后续生成都会基于此偏差,产生误导性延续,导致不可靠监督并浪费计算资源。我们发现失败前缀上师生延续存在不对称性,即教师倾向于重新引导,而学生继续沿原方向。在中继在线策略蒸馏(Relay-OPD)中,我们将此转化为无标签交接触发器。训练时,Relay-OPD通过让教师在检测到的触发点短暂接管以生成教师轨迹段,之后学生恢复并在生成的轨迹上优化。有限的中继预算将干预集中在关键早期位置,同时限制与学生策略的偏离。在八个数学推理基准测试中,使用Qwen3-4B-Instruct-2507教师和Qwen3-0.6B/1.7B-Non-Thinking学生,Relay-OPD在每个基准测试中都取得了最佳或次佳结果,平均比标准OPD在1.7B模型上提高了5.73%,比最强基线FastOPD提高了1.49%,在0.6B模型上也有持续提升,训练轨迹长度减少了50%以上。

英文摘要

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baseline FastOPD by +1.49% on average for 1.7B, with consistent gains at 0.6B. Training trajectory length is reduced by over 50%.

CommentsProject Page: https://zju-real.github.io/Relay-OPD Code: https://github.com/zju-real/Relay-OPD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑