发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出知识蒸馏的分布视角,将教师表示为多温度视角族,形式化设计空间并证明相关结果,在Pythia模型实验中得出三条经验规律,发现最佳KD损失取决于上限间隙Γ。
AI 中文摘要
令牌级知识蒸馏(KD)会匹配每个位置的两个条件分布,但标准目标却逐点比较它们:Kullback-Leibler梯度无法察觉哪个错误令牌获得了概率质量。我们提出一种分布视角,其中教师模型并非由单一的软化输出表示,而是由一系列多温度视角构成——即其logits退火路径的边缘分布;学生模型则在基于嵌入的地面代价下,针对这些视角的几何感知聚合进行训练。我们将由此产生的设计空间形式化(包括混合模型、对数线性池化、熵Wasserstein重心,以及中心和路径形式的去偏Sinkhorn散流旗舰方法),证明了一个精确的崩溃结果,即温度视角的对数线性池化等价于单一温度,并给出了多边际Schrödinger桥解读,该解读可产生可证伪的预测。在经过指令调优的Pythia模型对上,实验得出三条经验规律:(i)分散规律——多温度聚合的益处随视角的有效温度分散度单调增长,而非随其数量增长;(ii)分散视角解锁聚合算子——当基于传输的聚合开始优于平均时,重心与算术混合物恰好分离;(iii)由上限间隙Γ=PPL_SFT−PPL_T控制的双 regime 图景:当微调教师仅略优于监督学生时,温和的传输目标是最佳KD损失,但无KD能优于监督微调;而在实际上限处,排名反转——保真度-泛化相关性的符号翻转。我们认为“哪种蒸馏损失最佳”并非损失的固定属性,而是Γ的函数。
英文摘要
Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single softened output but by a family of multi-temperature views - marginals of the annealing path of its logits - and the student is trained against a geometry-aware aggregate of these views under an embedding-based ground cost. We formalize the resulting design space (mixtures, log-linear pooling, entropic Wasserstein barycenters, and a debiased Sinkhorn-divergence flagship in hub and path forms), prove an exact collapse result showing log-linear pooling of tempered views is equivalent to a single temperature, and give a multi-marginal Schrodinger-bridge reading that yields falsifiable predictions. On instruction-tuned Pythia pairs, experiments yield three empirical laws: (i) dispersion law - the benefit of multi-temperature aggregation grows monotonically with the effective temperature dispersion of the views, not with their number; (ii) dispersed views unlock the aggregation operator - the barycenter separates from the arithmetic mixture exactly when transport-based aggregation starts to beat averaging; and (iii) two-regime picture governed by the ceiling gap $Γ=\mathrm{PPL}_{\mathrm{SFT}}-\mathrm{PPL}_{T}$: when the fine-tuned teacher barely beats a supervised student the gentle transport objective is the best KD loss but no KD beats supervised fine-tuning, whereas at a real ceiling the ranking inverts - and the sign of the fidelity-generalization correlation flips. We argue that "which distillation loss is the best" is not a fixed property of the loss but a function of $Γ$.
Comments11 pages, 4 figures