训练中期的知识蒸馏更利于推理而非事实回忆
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
浏览论文内容
中文总结 AI 辅助
该研究发现训练中期的前向KD会减缓事实回忆但提升推理,提出Switch Distillation方法,在保留高推理和知识性能的同时维持大部分事实回忆,优于现有蒸馏目标。
中文摘要 AI 辅助
基于Logit的知识蒸馏(KD)通过更强教师模型的监督来训练更小的语言模型(LMs),但其益处是否在所有训练阶段都一致尚不清楚。通过控制实验,我们发现标准KD形式的前向Kullback-Leibler(KL)蒸馏,在训练中期(即对精选语料进行自监督学习的中间阶段),结合后训练教师模型时表现出根本不同的特性。令人惊讶的是,与标准的下一个token预测(NTP)相比,前向KD在预训练阶段同时提升了推理和事实回忆能力,但在训练中期却会减缓事实回忆的获取,尽管推理能力仍持续提升。我们将这种阶段依赖性归因于教师模型在不同数据域上的置信度不对称,以及学生模型不断变化的知识状态:教师在程序性数据上比知识密集型数据上更有信心,而学生在训练早期就会获取低熵的事实知识。为缓解这种不平衡,我们提出了Switch Distillation,这是一种简单的训练中期目标,利用教师预测熵作为轻量级路由信号,仅在教师有信心的token上进行蒸馏,否则回退到交叉熵。Switch Distillation在各种教师模型规模下始终优于现有蒸馏目标。与标准NTP相比,它实现了1.61-1.71倍的推理性能,以及1.13-1.19倍的知识和常识性能,同时保留了96.7-96.8%的事实回忆。至关重要的是,这些益处在后训练后仍然存在:Switch Distillation缩小了事实回忆差距,同时分别保持了1.25-1.32倍和1.13-1.20倍的推理和知识与常识增益。
英文摘要
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.
发表机构
- Meta AI
- University of Washington(华盛顿大学)
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。