arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

过犹不及——知识蒸馏何时会导致过拟合,以及如何避免

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

Irene Trigueros-Lorca, Leonardo Concepción, Christian Wagner, Isaac Triguero, Daniel Molina

arXiv 2608.23752首次发表:更新:

发表机构

Andalusian Research Institute in Data Science and Computational Intelligence; University of Granada(安达卢西亚数据科学与计算智能研究院; 格拉纳达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对知识蒸馏仅用于网络输出的局限,提出镜像教师网络分块的学生网络,发现细粒度数据稀缺场景用中间分块式蒸馏效果更好,可构建紧凑高效且精度不损失的模型。

AI 中文摘要

卷积神经网络规模的不断扩大催生了越来越庞大且昂贵的模型,知识蒸馏(Knowledge Distillation, KD)通过将大型网络(教师网络)的知识迁移至小型网络(学生网络)来解决这一问题,同时还能减少所需的训练数据。传统上,KD仅应用于网络的最终输出,但在网络中间层应用KD的行为却鲜少受到关注。这引发了一个问题:在特定条件下,比如细粒度数据集中常见的每类样本数量较少的情况,能为整个网络提供监督的中间分块式KD是否能带来优势?本研究提出一种基于简单、同质性分块的学生网络设计,该设计镜像教师网络的分块,在对应分块间进行知识蒸馏。在11个数据集上的实验表明,在经典数据集上,仅蒸馏最后一个分块就足够——且通常效果最佳;而在数据稀缺的细粒度场景中,中间监督能显著提升性能,仅增加一个额外的蒸馏点就能大幅缩小性能差距。我们还研究了应如何指导这种监督,探索了不同粒度的配置,并基于注意力图、中心核对齐(Centered Kernel Alignment)和Grad-CAM的可解释性分析,同时结合教师网络和学生网络的微调策略的影响展开研究。本研究表明,在恰当指导下的中间分块式蒸馏是构建紧凑、数据高效且不损失精度的模型的关键。

英文摘要

The growing size of Convolutional Neural Networks has led to increasingly large and costly models. Knowledge Distillation (KD) addresses this by transferring knowledge from a large network (teacher) to a small one (student), also reducing the training data required. KD is traditionally applied only at the network's final output. However, its behaviour when applied at intermediate network layers has received little attention. This raises the question of whether intermediate block-wise KD, which provides supervision throughout the network, could offer an advantage under specific conditions, such as few instances per class, which is common in fine-grained datasets. This work proposes a student design based on simple, homogeneous blocks mirroring those of the teacher, distilling knowledge between corresponding blocks. Across eleven datasets, we show that on classic datasets, distilling only the last block is sufficient -- and often best--, whereas fine-grained, data-scarce settings benefit substantially from intermediate supervision, with even a single additional distillation point narrowing the gap considerably. We further study how this supervision should be guided, exploring configurations of varying granularity and informed by an explainability analysis based on attention maps, Centered Kernel Alignment, and Grad-CAM, alongside the impact of teacher and student fine-tuning strategies. This work shows that intermediate block-wise distillation, guided appropriately, is key to building compact data-efficient models without sacrificing accuracy.

Comments19 pages, 7 images

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑