发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究诊断了同策略自蒸馏在数学推理语言模型中的效果,发现其仅在狭窄兼容性下有效,否则导致退化或崩溃,表明该算法敏感而非普遍可靠。
AI 中文摘要
同策略自蒸馏(OPSD)作为一种有前景的方法,日益受到关注,旨在提升语言模型的推理能力。在无需外部奖励或独立更强教师模型的情况下,具有特权信息的自教师模型能够在学生轨迹上提供密集信号。然而,其在语言推理中的行为仍不清楚,报告的结果从适度提升到行为崩溃不等。在本工作中,我们对参数规模从0.6B到8B的模型进行了数学推理上的同策略自蒸馏诊断。我们进行了受控实验和词元级分析,以深入探究同策略自蒸馏。我们指出,教师信号由推理模式对齐和完整教师前缀所塑造,而非仅由特权语义决定。同策略自蒸馏仅在狭窄的兼容性范围内提升推理能力。否则,它会导致无效的长度增长、稳定的性能退化或行为崩溃。词元级分析表明,教师信号不稳定,且不能预测下游性能。基于这些结果,我们认为同策略自蒸馏是一种敏感的算法,而非普遍可靠的推理改进后训练方法。
英文摘要
On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without external rewards nor a separate stronger teacher, the self-teacher with privileged information could provide dense signals on student's trajectories. However, its behavior in language reasoning remains unclear, with reported outcomes ranging from modest gains to behavioral collapse. In this work, we diagnose OPSD for mathematical reasoning across models spanning 0.6B--8B parameters. We conduct controlled experiments and token-level analyses to fully delve into OPSD. We point out that teacher's signal is shaped by reasoning-mode alignment and the complete teacher prefix, rather than by privileged semantics alone. OPSD improves reasoning only in narrow compatibility regimes. Otherwise, it produces ineffective length growth, stable degradation, or behavioral collapse. Token-level analysis shows that teacher's signal is not stable and does not predict downstream performance. Based on these results, we argue that OPSD is a sensitive algorithm rather than a generally reliable reasoning-improvement post-training method.