arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33455cs.AI

共享前缀所隐藏的问题:用于在线策略蒸馏的轨迹丢弃

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo

首次发表
浏览论文内容

中文总结 AI 辅助

针对在线策略蒸馏中共享前缀导致的监督衰减问题,提出轨迹丢弃方法,通过随机丢弃学生推理轨迹加强词元级监督,提升多规模模型对在数学推理基准上的性能。

中文摘要 AI 辅助

在线策略蒸馏(OPD)在由学生模型自身生成的轨迹上训练学生模型,并使用来自更强教师模型的密集词元级反馈。由于每次更新都基于学生模型已经生成的推理前缀,该前缀也决定了教师反馈转化为学习效果的效率。我们发现,共享前缀可能导致词元级更新变弱,我们将这一现象称为前缀引起的监督衰减(PISA)。这种衰减在两种常见情况下出现:(i)即使教师不同意,高学生置信度也可能削弱纠正梯度;(ii)依赖早期推理的词元可能接收到与简单局部延续同样弱的学习信号。为解决此问题,我们提出轨迹丢弃(Trajectory Dropout),一种简单的训练时干预方法,用于暴露这些被削弱的信号。学生模型首先执行标准的全上下文展开以生成完整轨迹。在训练期间,我们随机丢弃学生推理轨迹的特定比例,而教师模型继续观察完整轨迹以进行词元级监督。这种干预加强了对过度自信预测的纠正,并在前缀敏感位置引入额外监督。轨迹丢弃在不同规模的教师-学生模型对和六个数学推理基准上持续提高了平均性能,同时在两个域外基准上也取得了增益。它还可以灵活集成到现有的OPD变体中,且计算开销可忽略不计,进一步提升其性能。这些结果表明,轨迹丢弃提供了一种简单机制,可在不同模型规模和OPD目标下加强词元级监督。

英文摘要

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.

发表机构

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑