arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

蒸馏中的散度控制熵

Divergence controls entropy in distillation

Nicolas Zucchet, Scott W. Linderman

arXiv 2610.03529首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从熵视角研究蒸馏,证明前向KL膨胀学生熵,反向KL收缩熵,散度充当隐式熵正则化器,在自蒸馏中最佳超参数需补偿熵缩小。

AI 中文摘要

蒸馏已成为大型语言模型训练的核心原语,但其性质尚未被充分理解。我们采取熵的视角,研究学生模型的熵如何依赖于定义蒸馏目标的数据和散度。我们证明前向KL散度会使学生模型的熵膨胀至高于教师模型的熵。由于交叉熵训练是特殊情况,这产生了一个恒等式,我们在预训练和监督微调中定量验证了该恒等式。其他散度没有这样的保证:反向KL散度会使熵缩小,直到学生与教师之间的差距过大,而在这两者之间插值会在训练早期平滑地改变熵,但在收敛时则突然改变。在策略蒸馏的较低熵来自token级别的反向KL散度,而非在策略采样。因此,散度充当隐式熵正则化器,其作用在自蒸馏中最为清晰:当基于特权信息的条件作用使熵缩小,最佳工作的散度超参数是那些补偿这一效应的超参数。

英文摘要

Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑