用于因果充分可解释性的空间注意力噪声掩码
Spatial Attention Noise Masking for Causally Sufficient Interpretability
- Clemson University(克莱姆森大学)
- Binghamton University(宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出空间注意力噪声掩码框架,结合UNet风格掩码生成器与Resnet18编码器等,实现计算机视觉模型的因果充分可解释性,在五项分类任务中兼具掩码忠实性、分类性能与分布偏移鲁棒性。
AI中文摘要:
我们提出了一种用于计算机视觉模型可解释性的新型因果方法,该方法在分类前动态掩码输入图像。深度学习预测的可解释性在医学成像、安全和自动驾驶等高风险领域至关重要。大多数可解释性方法被动应用于已训练好的模型,通常产生的是相关性而非因果性解释。现有因果可解释性方法局限于事后分析,削弱了因果主张;此外,现有主动方法普遍缺乏明确将责任分配给输入特征的解释。本研究提出了一种空间注意力噪声掩码框架,可提供对预测充分特征的因果解释。该框架包含:1)UNet风格的掩码生成器;2)Resnet18编码器和线性分类器,用于对输入图像的掩码和未掩码版本进行分类。生成的掩码被正则化为稀疏且空间平滑,同时掩码后的图像嵌入被约束为与对应未掩码图像的嵌入保持一致。所得掩码可解释为特征归因图,其性能与相关可解释性方法相当,同时还能为模型预测提供强有力的因果解释。定量评估表明,该掩码具有忠实性,在五项分类任务中,尽管对图像信息进行了大量掩码,分类性能仍接近基线,且对背景交换、自然对抗样本等分布偏移具有鲁棒性。定性比较进一步表明,该掩码的行为与最先进的特征归因方法相比具有竞争力。
英文摘要:
We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging, security, and autonomous driving. Most interpretability methods are applied passively to already trained models, which typically result in correlational rather than causal explanations. Existing causal interpretability methods are limited to post hoc analysis, weakening the causal claims. Additionally, existing active methods generally lack explanations that explicitly assign responsibility to input features. This work proposes a spatial attention noise masking framework that provides causal explanations about the features sufficient for the prediction. The proposed framework consists of: 1) a UNet-style mask generator, and 2) a Resnet18 encoder and linear classifier that classifies both masked and unmasked versions of an input image. The generated masks are regularized to be sparse and spatially smooth, while masked image embeddings are constrained to remain consistent with embeddings from the corresponding unmasked images. The resulting masks can be interpreted as feature attribution maps that are competitive with related interpretability methods while additionally providing strong causal explanations of model predictions. Quantitative evaluations demonstrate mask faithfulness, near-baseline classification performance across five classification tasks despite substantial masking of image information, and robustness to distribution shifts such as background swapping and natural adversarial examples. Qualitative comparisons further demonstrate mask behavior and competitive interpretability relative to state-of-the-art feature attribution methods.