发表机构
Wuhan University of Technology; Tongji University(武汉理工大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出自适应互惠知识蒸馏(AR-KD),通过匹配教师类别相关性矩阵与学生关系表示来简化教师输出,缓解暗知识崩溃,在CIFAR-100和ImageNet-1k上显著提升学生性能。
AI 中文摘要
知识蒸馏旨在通过从更大、更强的教师模型中转移知识来提高轻量级学生模型的性能。然而,教师模型与学生模型之间巨大的规模差距常常阻碍有效的知识转移。大多数现有方法采用静态的、单向的教师到学生蒸馏范式,这忽视了学生学习的动态性,并且未能对困难样本提供有针对性的指导。在本文中,我们提出了自适应互惠知识蒸馏(AR-KD),一种通过简化教师输出分布来改进知识转移的新方法。具体而言,AR-KD通过将教师的类别相关性矩阵与学生模型的关联表示进行匹配,对教师进行互惠适配,从而重塑教师的预测结构以更好地适应学生的能力。这种关系对齐缓解了由过度自信的教师引起的类间暗知识崩溃,使学生能够从更丰富且更兼容的监督信号中学习。我们在CIFAR-100和ImageNet-1k分类数据集上评估了AR-KD,其性能优于最先进的知识蒸馏基线。具体而言,AR-KD在同构和异构设置中均提升了学生性能:学生准确率最高提升7.13%,平均比vanilla KD高出1.42%至4.15%,并且与其他先进方法集成时还有进一步提升。我们的代码可在该https URL获取。
英文摘要
Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher models. However, a substantial size gap between teacher and student models often impedes effective knowledge transfer. Most existing approaches adopt a static, one-way teacher-to-student distillation paradigm, which overlooks the dynamic nature of student learning and fails to provide targeted guidance on hard samples. In this paper, we propose adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer by simplifying the teacher's output distribution. Specifically, AR-KD performs reciprocal adaptation on the teacher by matching its class correlation matrix to the student's relational representation, which reshapes the teacher's prediction structure to better suit the student's capacity. This relational alignment mitigates the collapse of inter-class dark knowledge caused by overconfident teachers, enabling the student to learn from richer and more compatible supervisory signals. We evaluate AR-KD on CIFAR-100 and ImageNet-1k classification datasets, where it outperforms state-of-the-art knowledge distillation baselines. Specifically, AR-KD improves student performance across homogeneous and heterogeneous setups: up to 7.13% accuracy gain for students, 1.42% to 4.15% higher than vanilla KD on average, and further improvements when integrated with other advanced methods. Our code is available at https://anonymous.4open.science/r/ARKD/.
CommentsAccepted by 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)