发表机构
School of Computer Science and Engineering, Sun Yat-sen University; School of Science and Technology, Hong Kong Metropolitan University(中山大学计算机科学与工程学院; 香港都会大学科技学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对知识蒸馏难以平衡忠实模仿与鲁棒生成的问题,提出自适应熵蒸馏(AED)方法,通过分解反向KL目标函数并利用教师熵动态校准模仿强度,在指令遵循与数学推理基准上表现更优。
AI 中文摘要
知识蒸馏(KD)被广泛用于将大语言模型(LLM)的能力迁移到更小的学生模型,但现有目标函数往往难以在忠实模仿与鲁棒生成之间取得平衡。尤其值得注意的是,现有方法主要结合正向KL(FKL)与反向KL(RKL),却忽略了RKL本身提供了一种调整学生模型模仿强度的机制。受此启发,我们重新研究了在线策略反向Kullback-Leibler(RKL)蒸馏,并将其目标函数分解为教师拟合项与学生熵项,未引入显式的FKL分支。我们从理论上证明, token级最优学生分布对应教师分布的 tempered 变体,其中自适应权重控制着模式寻求与不确定性保留之间的权衡。基于这一见解,我们提出了自适应熵蒸馏(Adaptive Entropy Distillation,AED),它利用教师模型的熵动态校准 token 级模仿强度。在指令遵循与数学推理基准上的实验表明,AED实现了更优的整体性能,且通常能提升教师与学生模型的分布及熵对齐度。
英文摘要
Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faithful imitation and robust generation. In particular, existing methods mainly combine FKL and RKL, overlooking that RKL itself provides a mechanism for adjusting the student's imitation strength. Motivated by this, we revisit on-policy Reverse Kullback-Leibler (RKL) distillation and decompose its objective into a teacher-fitting term and a student-entropy term, without introducing an explicit FKL branch. We show theoretically that the token-level optimal student distribution corresponds to a tempered variant of the teacher distribution, where the adaptive weight controls the trade-off between mode-seeking and uncertainty preservation. Guided by this insight, we propose \textbf{Adaptive Entropy Distillation (AED)}, which uses the teacher's entropy to dynamically calibrate token-level imitation strength. Experiments on instruction-following and mathematical reasoning benchmarks demonstrate that AED achieves superior overall performance and generally improves teacher--student distributional and entropy alignment.