arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07803cs.AI

理解模型剪枝对医学影像中长尾遗忘与解释可靠性的影响

Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging

Nazish Khalid, Tausifa Jan Saleem, Amal Saqib, Donald C. Wunsch, Mohammad Yaqub

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统评估了模型剪枝对长尾医学影像的影响,发现低频类性能退化更早更大,解释可靠性主要受剪枝策略影响,建议评估应超越总体性能并兼顾类别与解释感知。

中文摘要 AI 辅助

模型剪枝被广泛用于压缩深度神经网络,以最小的总体性能影响减少内存和计算需求。然而,其对模型行为的影响仍知之甚少,尤其是在长尾医学数据集中,罕见但临床重要的病症代表性不足。此外,尚不清楚剪枝后的模型是否保留其预测的可靠解释。为弥补这一空白,我们提出了一项关于模型剪枝下长尾遗忘和解释可靠性的系统性研究。在两个长尾医学影像数据集、两种CNN架构、四种剪枝方法以及高达95%的稀疏度水平下,我们评估了预测性能、解释稳定性和解释忠实性。我们的结果表明,预测性能表现出强烈的频率依赖性趋势,低频类通常比高频类更早且更大幅度地退化。相比之下,解释稳定性和忠实性主要受剪枝策略影响,基于梯度的方法在激进压缩下能更有效地保持解释可靠性。定性和机制性分析进一步表明,解释退化主要与类判别梯度的崩溃相关,而非特征激活的消失。这些发现表明,模型压缩的评估应超越总体性能。引入类别感知和解释感知的评估揭示了原本隐藏的失败模式,而适度的稀疏度水平在压缩、预测性能和解释可靠性之间提供了实用的平衡。

英文摘要

Model pruning is widely used to compress deep neural networks, reducing memory and computational requirements with minimal impact on aggregate performance. However, its effect on model behavior remains poorly understood, particularly for long-tailed medical datasets where rare but clinically important conditions are underrepresented. Furthermore, it remains unclear whether pruned models preserve reliable explanations of their predictions. To address this gap, we present a systematic study of long-tail forgetting and explanation reliability under model pruning. Across two long-tailed medical imaging datasets, two CNN architectures, four pruning methods, and sparsity levels up to 95\%, we evaluate predictive performance, explanation stability, and explanation faithfulness. Our results show that predictive performance exhibits a strong frequency-dependent trend, with lower-frequency classes generally experiencing earlier and larger degradation than higher-frequency classes. In contrast, explanation stability and faithfulness are influenced primarily by the pruning strategy, with gradient-informed methods preserving explanation reliability more effectively under aggressive compression. Qualitative and mechanistic analyses further indicate that explanation degradation is primarily associated with the collapse of class-discriminative gradients rather than the disappearance of feature activations. These findings suggest that model compression should be evaluated beyond aggregate performance. Incorporating class-aware and explanation-aware evaluation reveals failure modes that would otherwise remain hidden, while moderate sparsity levels provide a practical balance between compression, predictive performance, and explanation reliability.

发表机构

  • Missouri University of Science and Technology(密苏里科技大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

↑