发表机构
University of Illinois at Chicago; Visa Research; Case Western Reserve University; University of Florida; The Ohio State University(伊利诺伊大学芝加哥分校; Visa研究院; 凯斯西储大学; 佛罗里达大学; 俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LastOPD通过在最后层状态仅短暂应用潜在监督并交叉淡入词元级在线策略蒸馏,解决了潜在在线策略蒸馏中的后期崩溃问题,显著提升数学推理性能。
AI 中文摘要
在线策略蒸馏(OPD)根据学生模型自己生成的响应进行纠正,但其信号是教师的下一词元分布:它告诉学生教师说了什么,却遗漏了教师如何思考。潜在监督通过将学生的潜在状态与教师的潜在状态对齐来弥补这一缺失部分。近期方法如OPRD将这种信号引入在线策略蒸馏。然而,在将Qwen3-4B和Qwen3-8B蒸馏到Qwen3-1.7B-Base时,我们观察到这种方案的两种失败。早期增益,后期崩溃:仅使用潜在监督在10步内将MATH-500准确率从25提升到46,但后续训练使性能降至11且无法恢复。更好的对齐,更差的行为:尽管对齐指标在整个崩溃过程中稳步提升,但最对齐的模型结果却是性能最差的。进一步分析表明,潜在信号的应用方式存在不匹配:按深度配对的层在两个模型中扮演不同角色,因此持续对齐可能将学生拉向其无法理解的教师状态。为解决此问题,我们提出LastOPD,它仅在最后一层状态(两个语言模型头共同读取的接口)应用潜在信号,并且仅在10步交叉淡入到词元级OPD期间使用。这保留了潜在信号的有用部分,并在崩溃发生前将学生交给词元级监督。大量实验表明,LastOPD在MATH-500上比仅词元级OPD分别提高5.55和4.02个百分点(使用4B和8B教师),在大多数保留数据集上领先,并且在大约一半的步数内达到仅词元级OPD的最终分数。代码可在该https URL获取。
英文摘要
On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery. Better alignment, worse behavior: although the alignment metric steadily improves throughout this collapse, the most aligned model turns out to be the worst performing. Further analysis suggests a mismatch in how the latent signal is applied: layers paired by depth play different roles in the two models, so continued alignment may pull the student toward teacher states it cannot understand. To address this, we propose LastOPD, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD. This keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in. Extensive experiments show that LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers, leads on most held-out datasets, and reaches the final score of token-only OPD in about half the steps. Code is available at https://github.com/Muyiiiii/LastOPD.