AI 中文总结
本文提出全面框架评估多种XAI方法在多数据集和模型上的可解释性,侧重保真度、简单性和稳定性,利用基准实验构建知识库以估计未见数据集和模型的可解释性分数,为评估比较XAI方法提供工具,助力可信AI系统发展。
AI 中文摘要
在本文中,我们提出了一个全面框架,用于评估诸如LIME和SHAP等各种可解释人工智能(XAI)方法在多个数据集和机器学习模型上的可解释性,最终目标是创建一个统一的多维可解释性分数。我们的方法侧重于可解释性的三个关键方面:保真度、简单性和稳定性。我们利用基准实验系统地评估这些方面,并利用获得的见解构建离线知识库。该知识库捕获每个注册模型的可解释性分数,并作为上下文相关可解释性评估的宝贵资源。通过分析人工智能模型、数据集和XAI方法的互补特征和元数据,该知识库将能够估计以前未见的数据集和模型的可解释性分数。保真度、简单性和稳定性等属性可能会因数据集、基础模型和最终用户的领域专业知识而有很大差异。我们通过将框架应用于三个开源数据集来展示我们的框架,并讨论所得结果与数据集特征相关的含义。我们的工作通过提供一个强大且通用的工具来评估和比较各种XAI方法的可解释性,为XAI领域的发展做出了贡献,最终支持更透明和可信的人工智能系统的开发。
英文摘要
In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score. Our methodology focuses on three key aspects of explainability: fidelity, simplicity, and stability. We leverage benchmarking experiments to systematically evaluate these aspects and use the insights gained to construct an offline knowledge base. This knowledge base captures the explainability scores for each registered model and serves as a valuable resource for context-dependent evaluation of explainability. By analyzing the complementary characteristics and metadata of AI models, datasets, and XAI methods, the knowledge base will enable the estimation of explainability scores for previously unseen datasets and models. Properties like fidelity, simplicity, and stability may vary significantly based on the dataset, underlying model, and domain expertise of the end user. We demonstrate our framework by applying it to three open-source datasets, discussing the implications of the obtained results in relation to the characteristics of the datasets. Our work contributes to the growing field of XAI by providing a robust and versatile tool for evaluating and comparing the explainability of various XAI methods, ultimately supporting the development of more transparent and trustworthy AI systems.
DOI:10.1109/DCOSS-IoT58021.2023.00084