发表机构
Universitat de les Illes Balears; Uiniversity of Pisa; ISTI-CNR(巴利阿里群岛大学; 比萨大学; 意大利国家研究委员会信息科学与技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于合成地面真值的框架,通过受控干预生成数据集,以评估九种XAI方法,揭示当前技术局限并强调干预式基准的重要性。
AI 中文摘要
评估可解释人工智能(XAI)方法是一项具有挑战性的任务,原因在于缺乏可靠的评估程序,尤其是缺少地面真值解释。在现有文献中,评估方法通常通过衡量解释相对于黑盒模型预测的保真度来评估解释。然而,此类评估策略仅量化了解释再现模型输出的程度,并未确保解释正确反映底层决策过程。因此,不同的解释可能获得相似的保真度分数,同时却提供不一致或误导性的模型行为解释。在本文中,我们提出了一种基于合成地面真值的XAI方法评估框架。所提出的方法依赖于受控干预来生成合成数据集,其中输入组件的重要性可通过设计确定。这使得能够构建与所分析模型行为直接对齐的地面真值解释。该框架在三个数据领域实例化,即二值图像、表格数据和时间序列,从而能够在异构设置中对解释方法进行全面评估。通过评估九种广泛使用的XAI方法获得的实验结果表明,当前技术存在显著局限性,并强调了基于合成的、干预式基准对于可靠评估解释质量的重要性。
英文摘要
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the explanation correctly reflects the underlying decision process. As a consequence, different explanations may achieve similar fidelity scores while providing inconsistent or misleading interpretations of the model behavior. In this paper, we propose a framework for the evaluation of XAI methods based on synthetic ground truth. The proposed approach relies on controlled interventions to generate synthetic datasets in which the importance of input components can be determined by design. This enables the construction of ground truth explanations that are directly aligned with the behavior of the model under analysis. The framework is instantiated across three data domains, namely binary images, tabular data, and time series, allowing a comprehensive assessment of explanation methods in heterogeneous settings. Experimental results obtained by evaluating nine widely used XAI methods show significant limitations in current techniques and highlight the importance of synthetic, intervention-based benchmarks for a reliable assessment of explanation quality.