发表机构
Alibaba Cloud Computing(阿里云计算)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究跨分词器同策略蒸馏,发现严格对齐位置的紧凑监督比扩大覆盖更有效,提出优先监督可靠性以提升性能。
AI 中文摘要
同策略蒸馏(OPD)利用教师反馈,在学生自身生成的内容上训练学生。当使用不同的分词器时,比较教师和学生的预测需要在序列和词汇两个层面进行对齐。本文考察了扩大这种对齐覆盖范围是否有助于提升学习效果。在数学推理和代码生成任务上,针对三组异构的教师-学生组合,尽管词汇表存在显著不匹配,严格的1:1分组已经覆盖了学生生成的大部分词元。在蒸馏前从学生采样的响应中,共享词汇在严格对齐的位置上平均保留了教师和学生几乎所有的概率质量。将反向KL散度限制在严格位置每个位置上学生选定的共享词汇前16个子集,其准确率与使用完整共享词汇的OPD相当,并且优于所评估的跨分词器基线方法。在错配分组中,对跨度对数概率添加均方误差监督可提供完整的监督覆盖,但会降低准确率。在仅使用严格损失训练的检查点上,跨度梯度与严格梯度的方向一致性较弱或为负,且其幅度相对于严格梯度有所增长。这些诊断可能有助于解释添加跨度监督导致的准确率下降。我们的研究结果促使从最大化对齐覆盖转向优先考虑监督可靠性:在严格位置上的紧凑监督可能比提供弱对齐或冲突训练信号的更广泛覆盖更有效。
英文摘要
On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.