发表机构
ETH Zurich; Max Planck Institute for Intelligent Systems(苏黎世联邦理工学院; 马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出对比自蒸馏,结合正确解教师的吸引与错误解教师的排斥,以分离正确性与行为信号,提升推理性能并保持响应长度稳定。
AI 中文摘要
同策略自蒸馏通过将模型置于特权信息条件下,并将由此产生的教师分布蒸馏回模型,提供密集的、词元级别的监督。然而,特权信息不仅可能改变教师所知道的内容,还可能改变其行为方式,从而将正确性相关的学习信号与意外的行为偏移纠缠在一起。我们通过对比吸引型自蒸馏(将模型移向特权教师)与排斥型自蒸馏(将模型移离特权教师)来研究推理任务中的这一效应。我们发现,这两种目标都能引发强烈且相反的行为偏移:吸引抑制探索性推理,并促使生成更短、更自信的响应;而排斥则增加响应长度,可能触发模型潜在思考模式的意外切换,并最终变得不稳定。受这些观察的启发,我们研究了对比自蒸馏,它将朝向正确解条件教师的吸引与远离错误解条件教师的排斥相结合。与先前将此类蒸馏信号与GRPO目标相结合的工作不同,我们单独隔离自蒸馏目标并研究其自身行为。我们发现,两位教师共享的行为偏移在很大程度上相互抵消,留下一个更直接反映正确性的词元级信号。在非思考、仅指令和已思考模型中,这种对比目标提高了推理性能,同时保持了稳定的响应长度。
英文摘要
On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change not only what the teacher knows, but also how it behaves, entangling correctness-relevant learning signals with unintended behavioral shifts. We study this effect in reasoning tasks by contrasting attractive self-distillation, which moves the model toward a privileged teacher, with repulsive self-distillation, which moves it away from a privileged teacher. We find that both objectives can induce strong and opposing behavioral shifts: attraction suppresses exploratory reasoning and promotes shorter, more confident responses, whereas repulsion increases response length, can trigger unintended switches into a model's latent thinking mode, and ultimately becomes unstable. Motivated by these observations, we study contrastive self-distillation, which combines attraction toward a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned teacher. In contrast to prior work that combines such distillation signals with a GRPO objective, we isolate the self-distillation objective and study its behavior on its own. We find that the shared behavioral shifts of the two teachers largely cancel, leaving a token-level signal that more directly reflects correctness. Across non-thinking, instruct-only, and already-thinking models, this contrastive objective improves reasoning performance while maintaining stable response lengths.