arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

忠实性评估的忠实性:乳腺MRI案例研究中的挑战与陷阱

Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study

Peachapong Poolpol, Henrik H. J. Detjen, Eike Petersen

arXiv 2609.25978首次发表:更新:

发表机构

Fraunhofer Institute for Digital Medicine MEVIS; Deggendorf Institute of Technology(弗劳恩霍夫数字医学研究所; 代根多夫应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过乳腺MRI分类案例,揭示显著性图忠实性评估中扰动策略与类别特定性两大挑战,并提出公平比较方法,发现Grad-CAM和Gradient Attention Rollout为最强类别特定方法。

AI 中文摘要

显著性图被广泛用于解释医学影像中的深度学习预测,然而视觉上合理的解释并不一定反映模型真实的决策过程,因此可能误导临床医生。我们使用基于Vision Transformer的乳腺MRI分类器(在ODELIA Breast MRI Challenge数据集上训练)研究这一问题,并评估多种显著性方法,包括Last-layer Attention、Attention Rollout、Grad-SAM、Gradient Attention Rollout、GMAR、Grad-CAM和HiResCAM。我们的研究强调了基于扰动的忠实性评估中两个常被忽视的挑战。首先,方法排名强烈依赖于扰动策略,在基于强度的扰动和基于Transformer的注意力掩蔽之间有所不同。其次,对显著性方法进行基准测试需要区分类别特定和类别无关的解释。为了进行公平比较,我们引入了基于梯度的方法的非类别特定变体,并分别评估这两种设置。在各种协议中,Grad-CAM和Gradient Attention Rollout始终是最强的类别特定方法,尽管它们的相对排名取决于评估设计。这些发现揭示了当前基于显著性的可解释性方法的重要局限性,并强调需要更稳健和标准化的评估框架,以实现可信赖的临床AI系统。

英文摘要

Saliency maps are widely used to explain deep learning predictions in medical imaging, yet visually plausible explanations do not necessarily reflect a model's true decision process and may therefore mislead clinicians. We investigate this problem using a Vision Transformer-based breast MRI classifier trained on the ODELIA Breast MRI Challenge dataset and evaluate multiple saliency methods, including Last-layer Attention, Attention Rollout, Grad-SAM, Gradient Attention Rollout, GMAR, Grad-CAM, and HiResCAM. Our study highlights two often-overlooked challenges in perturbation-based faithfulness evaluation. First, method rankings depend strongly on the perturbation strategy, varying across intensity-based perturbations and transformer-based attention masking. Second, benchmarking saliency methods requires distinguishing between class-specific and class-agnostic explanations. To enable fair comparisons, we introduce non-class-specific variants of gradient-based methods and evaluate both settings separately. Across protocols, Grad-CAM and Gradient Attention Rollout consistently emerged as the strongest class-specific methods, although their relative ranking depended on the evaluation design. These findings expose important limitations of current saliency-based explainability approaches and highlight the need for more robust and standardized evaluation frameworks for trustworthy clinical AI systems.

CommentsAccepted at MICCAI iMIMIC Workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑