通过其诱导的预测器来评估解释方法
Evaluating Explanation Methods by the Predictors They Induce
浏览论文内容
中文总结 AI 辅助
提出通过将解释转化为预测器并测试其复现能力来评估解释方法,证明特征依赖性决定方法优劣,SHAP在相关特征上领先。
中文摘要 AI 辅助
机器学习模型的解释通常通过难以比较的标准来评判。我们提出一个更简单的测试:如果一个解释确实描述了模型如何使用其特征,那么应该能够从该解释重建模型的预测。我们通过读取每个特征的效果并将它们相加,将每个解释转化为一个预测器,并衡量该预测器在未见数据上复现模型的程度。不进行任何拟合,因此得分反映了解释本身。该测试适用于任何可以写成特征函数的解释;我们在部分依赖图(PDP)、累积局部效应(ALE)、SHAP和LIME上进行了演示。我们证明,当特征独立时,对部分依赖曲线求和可以得到模型的最佳加性摘要,而当特征相关时,这一方法会失效。在13个真实数据集、9个合成设计和4个模型家族中,哪种方法得分最高完全取决于特征依赖性:在特征独立的情况下,SHAP略逊于PDP,正如理论所预测;在特征相关的真实数据上,SHAP领先。一些广泛使用的质量指标甚至更倾向于损坏的解释而非完整的解释。
英文摘要
Explanations of machine learning models are usually judged by criteria that are hard to compare. We propose a simpler test: if an explanation really describes how a model uses its features, it should be possible to rebuild the model's predictions from it. We turn each explanation into a predictor by reading each feature's effect and adding them up, and measure how well that predictor reproduces the model on unseen data. Nothing is fitted, so the score reflects the explanation itself. The test applies to any explanation that can be written as a function of the features; we demonstrate it on partial dependence plots (PDP), accumulated local effects (ALE), SHAP and LIME. We prove that summing partial dependence curves gives the best possible additive summary of a model when its features are independent, and that this fails when they are dependent. Across 13 real datasets and 9 synthetic designs and four model families, which method scores best depends entirely on feature dependence: where features are independent SHAP is slightly worse than PDP, exactly as the theory predicts; on dependent real data SHAP leads. Some widely used quality metrics even prefer a damaged explanation to an intact one.
发表机构
- Oslo Metropolitan University(奥斯陆城市大学)
- SimulaMet
- Oslo University Hospital(奥斯陆大学医院)
机构由 AI 辅助整理,请以论文原文为准。