发表机构
Barcelona Supercomputing Center(巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于分布的归因可分性框架,用于评估归因方法的稳定性及对比其排名鲁棒性,为归因方法稳定性评估提供补充准则。
AI 中文摘要
归因方法(AMs)为每个特征分配重要性得分,被广泛用于解释黑盒模型。然而,大多数方法因自身定义中的随机成分会产生可变的归因得分。本文提出一种基于分布的框架以捕捉归因得分的稳定性,该方法可理解排名归因向量中的可分性程度,并获得特征排名保持可靠的最大索引;我们还扩展该框架,以基于数据集上排名的鲁棒性对比不同AMs。通过实验,我们展示如何将该方法用于评估解释器稳定性,总体而言,该方法为评估AMs的稳定性提供了补充性准则。
英文摘要
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and obtain the largest index for which a feature ranking remains reliable. We further extend this framework to compare AMs based on the robustness of their rankings across a dataset. Through experiments, we demonstrate how to apply our method to evaluate explainer stability. Overall, our approach provides a complementary criterion for evaluating the stability of AMs.
CommentsAccepted at EXPLAINS 2026 Conference