发表机构
University of Luxembourg; Seafill Open-Source Community; Université Paris-Saclay(卢森堡大学; Seafill 开源社区; 巴黎-萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文分析小规模同策略蒸馏,发现其传递求解能力但不传递停止能力,学生越小差距越大,并提出区分答案标记、正确性与停止的诊断方法。
AI 中文摘要
同策略蒸馏是一种将推理能力传递给较小模型的常见方法,其中学生模型基于自身输出,从更强教师模型的反馈中学习。我们分析了该方法在小规模下所传递的内容,将Qwen3-8B蒸馏至Qwen3 4B、1.7B和0.6B的学生模型,分别在思考模式(先进行长推理,然后结束推理并给出答案)以及非思考模式(无独立推理阶段)下进行对比。长推理需要两种能力:求解问题以及判断何时已求解。我们发现,蒸馏传递了第一种能力,但在思考模式下并未传递第二种能力。求解能力在每个规模上均有所提升,直至达到两个上限,我们全面测量了两种模式及所有学生规模下的表现:学生单次尝试的成果不会超过其在训练前多次尝试所能达到的水平,且学生模型越小,其与教师模型的差距越大。停止能力是两种模式的分歧点。在非思考模式下,所有学生模型均能保持停止;在思考模式下,学生模型在训练早期便停止结束推理,且学生模型越小,该能力保留得越少:教师模型几乎只在学生已结束推理之处发出停止信号,因此蒸馏并未教会新的停止行为;它仅保留了学生模型已有的、落在正确答案上的停止行为,而弱学生模型此类停止行为很少。最小的学生模型往往能得出正确的数值,但并未将其确定为最终答案:它们要么很少标记该答案,要么标记后继续书写而越过它。综合来看,这些结果描述了小规模学生模型在同策略蒸馏下的行为,并提供了一种区分答案标记、正确性与停止行为的诊断方法。
英文摘要
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.
Comments22 pages, 13 figures