arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越完全模仿的学习:任务保持的知识蒸馏

Learning Beyond Full Imitation: Task-Preserving Knowledge Distillation

Qianfeng Yuan, Wenbing Tao

arXiv 2609.39338首次发表:更新:

发表机构

Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出任务保持知识蒸馏(TPKD),在保持标签梯度完整的同时最小修正条件梯度,以在完全模仿受阻时仍能有效传递知识,并在CIFAR-100和CLINC150上超越标准蒸馏。

AI 中文摘要

知识蒸馏通过鼓励学生模型匹配教师模型预测的类别概率来传递知识。这些概率不仅表达了对正确类别的置信度,还表达了错误备选类别之间的关系。然而,更紧密的模仿并不必然产生更好的学生模型。学生模型可能已经比教师模型更清晰地区分正确类别,因此进一步的模仿可能需要放弃其已获得的判别能力。我们的主要结果是在完全模仿和条件学习之间实现了精确的分离。当必须保持正确类别相对于每个备选类别的得分优势时,当且仅当学生模型分配给每个错误类别的概率不超过教师模型时,完全的教师到学生的KL最小化被阻碍。关键在于,教师模型在错误类别之间的相对概率仍然是完全可学习的。我们刻画了这种迁移的精确代价:正确类别对数几率的最小增加,以补偿最大的条件概率不匹配。因此,标签拟合和条件匹配可以在完全教师KL发散的情况下完成。这种分离促使了任务保持知识蒸馏(TPKD)的提出,该方法保持标签梯度不变,并最小程度地修正条件梯度,使其输出更新保持标签步骤相对于每个错误备选类别的增益。修正后的条件方向在相同步长下保留了原始一阶条件下降的一半以上,且具有紧界。对于固定的正条件目标和足够小的恒定输出步长,标签和条件误差会同时消失。实验从精确的头部更新追踪到普通网络训练。TPKD在CIFAR-100上达到88.05%的准确率,在CLINC150上达到93.81%,在三个随机种子上比标准蒸馏分别提高了0.47和0.35个百分点。

英文摘要

Knowledge distillation transfers knowledge by encouraging a student to match a teacher's predicted class probabilities. These probabilities express not only confidence in the correct class, but also relations among incorrect alternatives. Yet closer imitation does not necessarily yield a better student. A student may already distinguish the correct class more sharply than its teacher, so further imitation can require giving back discrimination it has acquired. Our main result is an exact separation between full imitation and conditional learning. When the correct class's score advantage over each alternative must be preserved, full teacher-to-student KL minimization is blocked exactly when the student assigns no more probability than the teacher to every incorrect class. Crucially, the teacher's relative probabilities among incorrect classes remain fully learnable. We characterize the exact price of this transfer: a minimum increase in correct-class log-odds that compensates for the largest conditional-probability mismatch. Label fitting and conditional matching can therefore be completed even as full teacher KL diverges. This separation motivates task-preserving knowledge distillation (TPKD), which keeps the label gradient intact and minimally corrects the conditional gradient so that its output update preserves the label step's gains against every incorrect alternative. The corrected conditional direction retains more than half of the original first-order conditional descent at the same step size, with a tight bound. For a fixed positive conditional target and sufficiently small constant output steps, label and conditional errors vanish together. Experiments trace this learning from exact head updates to ordinary network training. TPKD reaches 88.05% accuracy on CIFAR-100 and 93.81% on CLINC150, improving over standard distillation by 0.47 and 0.35 percentage points across three seeds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑