arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迁移视觉解释:跨架构知识蒸馏如何影响模型可解释性

Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability

Aleks Czufarow, Ihor Babin

arXiv 2609.23561首次发表:更新:

发表机构

XIV High School of Stanislaw Staszic; Ukrainian Catholic University(斯坦尼斯瓦夫·斯塔希茨第十四中学; 乌克兰天主教大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过跨架构知识蒸馏实验,发现蒸馏主要传递空间关注区域而非像素级归因,可解释性对软标签权重敏感,且受学生模型结构偏差限制。

AI 中文摘要

在资源受限的环境中部署高效神经网络至关重要,然而紧凑模型往往牺牲可解释性——这在自动驾驶和医学等安全关键领域尤为关键。本研究探讨知识蒸馏是否将大型教师网络的空间特征归因迁移至紧凑学生网络。为评估蒸馏方案对可解释性的影响,我们在ImageNet-1K上通过系统改变蒸馏温度和软标签损失权重,将ResNet-152教师蒸馏为ResNet-34学生,共配置五种方案。模型评估采用top-1准确率以及两个可解释性指标:相关性质量准确率和相关性排名准确率。这些指标通过Grad-CAM热图计算,并以真实目标掩码为基准。结果显示,top-1准确率范围为71.6%至74.0%。对于Grad-CAM,RMA范围为7.7%至9.7%,RRA范围为7.3%至10.1%;对于Guided Grad-CAM,RMA范围为16.1%至18.6%,RRA范围为15.9%至21.5%。可解释性对软标签权重的敏感性远高于温度:将学生锚定于硬标签可同时保持准确率和粗略定位,而大幅加权教师则使两者均下降。然而,细粒度归因在所有测试配置中均低于未蒸馏基线,表明logit蒸馏更易传递模型关注的区域,而非该关注的像素级结构。我们评估了12种卷积和基于Transformer模型的跨架构组合,揭示细粒度空间推理的继承从根本上受限于学生模型的内在结构偏差。据我们所知,这是该可解释性感知评估框架(先前用于神经网络剪枝)首次应用于知识蒸馏。

英文摘要

Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical in safety-critical domains such as autonomous driving and medicine. This study investigates whether Knowledge Distillation transfers the spatial feature attribution of a large teacher network to a compact student. To assess the influence of the KD scheme on interpretability, we distill a ResNet-152 teacher into a ResNet-34 student on ImageNet-1K across five configurations by systematically varying the distillation temperature and soft-label loss weight. Models are evaluated on top-1 accuracy, along with two interpretability metrics: Relevance Mass Accuracy and Relevance Rank Accuracy. These metrics are computed via Grad-CAM heatmaps benchmarked against ground-truth object masks. Our results show that top-1 accuracy ranges from 71.6% to 74.0%. For Grad-CAM, RMA ranges from 7.7% to 9.7% and RRA from 7.3% to 10.1%; for Guided Grad-CAM, RMA ranges from 16.1% to 18.6% and RRA from 15.9% to 21.5%. Interpretability proves far more sensitive to the soft-label weight than to the temperature: keeping the student anchored to hard labels preserves both accuracy and coarse localization, whereas weighting the teacher heavily degrades both. Fine-grained attribution, however, fell below the undistilled baseline in every configuration tested, indicating that logit distillation transmits where a model attends more readily than the pixel-level structure of that attention. We evaluate 12 cross-architecture combinations of convolutional and transformer-based models, revealing that the inheritance of fine-grained spatial reasoning is fundamentally bottlenecked by the student's intrinsic structural biases. To our knowledge, this is the first application of this interpretability-aware evaluation framework - previously used for neural network pruning - to KD.

Comments21 pages, 4 figures, 2 tables. Submitted to AJOSR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑