发表机构
Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有Top-k在线策略蒸馏丢弃长尾概率信号的问题,提出TA-OPD方法,通过引入携带长尾概率的长尾词优化,在常见基准上使Avg@8最高提升8.05个百分点。
AI 中文摘要
在线策略蒸馏(OPD)已成为在语言模型间迁移知识的有效范式,其中学生模型被训练以沿自身轨迹将下一个词的分布与教师模型对齐。为以可处理的成本提供密集监督,许多工作最小化学生模型与教师模型在教师模型的Top-k个词上归一化分布之间的反向KL散度。然而,该归一化目标丢弃了关于长尾概率的信息,即教师模型Top-k个词之外的总概率。结果,优化过程可能会稳步增加学生模型的长尾概率和熵,经验上降低下游任务的准确率。为解决此问题,我们提出感知长尾的Top-k在线策略蒸馏(TA-OPD),这是一种恢复缺失长尾概率信号的新型蒸馏方法。具体而言,TA-OPD最小化Top-k个词上的反向KL散度,再加上携带长尾概率的长尾词。实际上,TA-OPD能更好地将学生模型的下一个词分布与教师模型对齐,防止Top-k归一化导致的长尾概率和熵的增加。大量实验证明了TA-OPD的优越性,在常见基准上使Avg@8提升了最高8.05个百分点。我们的代码可在该https URL获取。
英文摘要
On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories. To provide dense supervision at tractable cost, many works minimize the reverse Kullback-Leibler (KL) divergence between the student and teacher's normalized distributions over the teacher's top-$k$ tokens. However, this normalized objective discards the information about tail probability: the total probability outside the teacher's top-$k$ tokens. As a result, the optimization can steadily increase the student's tail probability and entropy, empirically degrading downstream accuracy. To address this issue, we propose Tail-Aware Top-$k$ OPD (\textbf{TA-OPD}), a novel distillation method that restores the missing tail probability signal. In particular, TA-OPD minimizes the reverse KL divergence over the top-$k$ tokens plus a tail token that carries the tail probability. In effect, TA-OPD better aligns the student's next-token distribution with the teacher's, preventing the increase in tail probability and entropy caused by top-$k$ normalization. Extensive experiments demonstrate the superiority of TA-OPD, improving Avg@8 by up to 8.05 points on common benchmarks. Our code is available at https://github.com/HuipengHuang/TA-OPD.