arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

立场:解释稳定性是模型与方法对的属性,而非模型本身的属性

Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model

Kabilan Elangovan, Daniel Ting

arXiv 2607.16652首次发表:更新:

AI 中文总结

该论文指出关于解释稳定性的断言需经交叉方法验证,以胸部X光实验为例,不同模型在不同归因方法下稳定性排名反转,说明解释稳定性是模型与方法对的属性,基于解释的断言应多方法验证并明确归因算子。

AI 中文摘要

本立场文件认为,未经交叉方法验证,关于解释稳定性的断言在科学上是无效的。正如统计显著性需要指定检验统计量一样,稳定性应要么跨多个归因范式进行评估,要么明确限定于单一方法的计算目标。在胸部X光控制实验中,DenseNet201、ResNet50V2和InceptionV3的AUC值均高于99%,但其稳定性排名在不同归因方法间反转。LayerCAM将InceptionV3列为最稳定模型,IoU为0.777,而GradCAM++更青睐DenseNet201,且将InceptionV3稳定性得分降低了17.3%。这些发现表明解释稳定性是模型与方法对的涌现属性,而非仅为模型的固有特征。因此我们认为基于解释的断言应通过多种归因方法进行验证,监管提交应明确指定所用的归因算子,以避免产生虚幻的安全保证。

英文摘要

This position paper argues that claims about explanation stability are scientifically invalid without cross method validation. Just as statistical significance requires the test statistic to be specified, stability should either be evaluated across multiple attribution paradigms or explicitly scoped to the computational objective of a single method. In controlled chest X ray experiments, DenseNet201, ResNet50V2, and InceptionV3 achieved AUC values above 99%, yet their stability rankings reversed across attribution methods. LayerCAM ranked InceptionV3 as the most stable model, with an IoU of 0.777, whereas GradCAM++ favored DenseNet201 and reduced InceptionV3 stability score by 17.3%. These findings demonstrate that explanation stability is an emergent property of the model method pair rather than an intrinsic characteristic of the model alone. We therefore argue that explanation based claims should be validated across multiple attribution methods and that regulatory submissions should explicitly specify the attribution operators used to avoid creating illusory safety assurances.

CommentsAccepted Position Paper at ICML 2026. https://openreview.net/forum?id=9r8sSaYhos

Journal refForty-third International Conference on Machine Learning Position Paper Track, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑