特权信息为同策略自蒸馏带来了什么?
What Does Privileged Information Add to On-Policy Self-Distillation?
浏览论文内容
中文总结 AI 辅助
本研究通过构建AMPLE-Math套件,隔离了同策略自蒸馏中特权参考信息的贡献,发现其价值在于促进跨模式迁移,而非揭示解决方案内容。
中文摘要 AI 辅助
同策略自蒸馏(OPSD)让语言模型从自身的冻结副本中学习,该副本能够看到答案或完整的解题过程。给予教师这一额外信息似乎为学生提供了更多可学习的内容,但相较于蒸馏本身,它究竟增加了多少价值?为隔离这一贡献,我们构建了AMPLE-Math,一个包含5,319道数学问题、具有六种共享同一答案的推理视角的可复用套件,并将每个视角与匹配的无参考蒸馏进行比较。在具备思考能力的教师监督直接响应生成的情况下,无参考蒸馏在思考能力评估下,无论是领域内还是外部基准上,都解释了Qwen3-1.7B的大部分改进。在Qwen中,额外参考益处的证据较为温和,对精炼解答最为显著,而完整轨迹在SmolLM3-3B第50步时增加了两个百分点。这些益处取决于学生是否接受训练。在同一检查点,将短直接响应生成替换为长思考能力生成,在两个模型家族中都使收益转为损失,而问题、参考和评估保持不变。Qwen中的教师配置文件和匹配的损失干预进一步表明,改变词元级监督可以使学生行为基本保持不变。综合这些发现表明,OPSD可以通过直接响应和思考能力推理共享的参数来改善对现有推理能力的访问。特权参考的价值在于它为这种跨模式迁移增加了什么,而不是它揭示了解决方案的多少。
英文摘要
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for an additional reference benefit is modest in Qwen, strongest for a polished solution, whereas complete traces add two percentage points in SmolLM3-3B at step 50. These benefits depend on the student being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families while the problems, references, and evaluation stay fixed. Teacher profiles and matched loss interventions in Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference. The value of a privileged reference is what it adds to this cross-mode transfer, not how much of the solution it reveals.
发表机构
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。