当数据更少时,较小的教师模型更优:重新思考知识蒸馏中的教师容量与数据选择
When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
- Pohang University of Science and Technology (POSTECH)(浦项科技大学(POSTECH))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究知识蒸馏中数据剪枝下教师容量与数据预算的关系,提出DVA方法,利用小教师进行难度过滤和关系容量最大化,无需训练动态即可优于免训练基线。
AI中文摘要:
数据剪枝降低了知识蒸馏(KD)的训练成本。然而,偏好的教师容量会随数据预算而变化:当可用训练数据有限时,较小的教师模型可以胜过较大的教师模型。理解驱动这一转变的因素不仅对教师选择很重要,而且对识别哪些样本对蒸馏有用也很关键。我们通过将教师监督分解为关系排序(类别的排名)和分数几何(类别概率的幅度和边界)来分析教师监督,并表明在低数据机制下,小教师优势不仅源于分数几何,还源于关系排序。除了理解教师容量外,我们的分析还揭示了有效子集的两个特性:样本应匹配适合可用预算的难度,并且它们的关系信号应多样化而非冗余。基于这些发现,我们提出了DVA(用于KD的难度和容量感知数据选择),一种免训练动态的方法,该方法使用小教师作为预算感知难度过滤和类条件关系容量最大化的代理。尽管不需要训练动态统计,我们的方法仍与基于训练动态的方法保持竞争力,同时始终优于免训练动态的基线方法。
英文摘要:
Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.