基于扰动的可解释人工智能中的分歧——合成孔径雷达图像洪水检测的扰动选择基准测试
On the Disagreement in Perturbation-based xAI -- Benchmarking Perturbation Choices for Flood Detection from SAR Images
浏览论文内容
中文总结 AI 辅助
研究基于扰动的xAI在洪水检测中不同扰动设置下相关性估计的变化,关注补丁几何形状和扰动类型两个关键参数,通过实验评估其一致性和忠实度,强调方法设置重要性,凸显审慎评估扰动选择的必要性。
中文摘要 AI 辅助
基于扰动的可解释人工智能(xAI)方法被广泛用于分析深度学习模型的行为和预测。通过改变输入区域并测量相对于原始图像的类别概率变化,它们分配相关性分数并生成反映每个区域对预测贡献的热图。然而,尽管表面简单,但基于扰动的方法对参数选择敏感。在这项工作中,我们关注扰动流程的两个关键参数,即补丁几何形状(包括扰动区域的大小和形状)和扰动类型(由替换方案定义)。基于合成孔径雷达图像洪水检测的用例,我们全面研究了在不同扰动设置下相关性估计如何变化。除了目视检查生成的相关性图之外,我们评估它们在不同扰动策略之间的一致性以及它们对模型推理的忠实度。我们展示了不同的扰动选择如何引导生成的相关性图,产生模糊甚至矛盾的解释。我们的发现强调了基于扰动的xAI中方法设置的重要性。它们强调了仔细检查和评估扰动选择并将其作为解释时不可或缺的一部分的必要性,以确保对解释和模型预测有稳健的理解。
英文摘要
Perturbation-based xAI methods are widely used to analyze the behavior and predictions of deep learning models. By altering input regions and measuring the resulting changes in class probabilities with respect to the original image, they assign relevance scores and generate heatmaps that reflect each region's contribution to the prediction. Despite their apparent simplicity, however, perturbation-based methods are sensitive to parameter choices. In this work, we focus on two key parameters of the perturbation pipeline, namely the patch geometry, including the size and shape of the perturbed regions, and the perturbation type, defined by the replacement scheme. Grounded in the use case of flood detection from Synthetic Aperture Radar imagery, we conduct a comprehensive investigation of how relevance estimation changes under different perturbation settings. Beyond visual inspection of the resulting relevance maps, we evaluate their consistency across perturbation strategies and their faithfulness to the model's reasoning. We demonstrate how different perturbation choices can steer the resulting relevance maps, yielding ambiguous and even contradictory explanations. Our findings emphasize the importance of methodological settings in perturbation-based xAI. They underscore the need to carefully inspect and evaluate perturbation choices and to treat them as an integral part when interpreting explanations, ensuring a robust understanding of both the explanations and model predictions.