arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CEL:综合反事实解释库与基准测试

CEL: Comprehensive Counterfactual Explanations Library and Benchmark

Oleksii Furman, Łukasz Lenkiewicz, Marcel Musiałek, Maciej Zięba

arXiv 2607.22045首次发表:更新:

发表机构

Wrocław University of Science and Technology(弗罗茨瓦夫科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对可解释人工智能中反事实解释方法评估难的问题,引入CEL统一库与基准测试,包含多数据集及方法实现,通过标准化设置全面比较多种方法,为反事实解释方法评估提供统一框架,促进未来方法发展。

AI 中文摘要

反事实解释是可解释人工智能(xAI)中的一种重要方法,能为改变模型预测提供可行指导。早期方法聚焦最小特征变化,近期工作纳入稀疏性等属性。但公平系统的评估仍具挑战,现有研究依赖不同数据划分等,限制了方法间的客观比较。为此,我们引入CEL,它是一个统一库与基准测试,包含18个不同规模和复杂度的数据集及14种反事实方法的实现。利用标准化设置,我们在不同数据集上对多种方法进行全面定量比较,评估协议包含多个互补指标。这是首个在统一可复现框架内系统评估反事实解释方法的综合基准测试,旨在提高可复现性、实现公平比较并为未来方法开发搭建平台。

英文摘要

Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minimal feature changes, recent work incorporates additional properties such as sparsity, actionability and plausibility. Despite this progress, fair and systematic evaluation remains challenging. Existing studies often rely on different data splits, predictive models, and evaluation metrics, which limits objective comparison across methods. To fill this gap, we introduce CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations designed to support consistent implementation and evaluation. CEL includes 18 datasets of varying size and complexity and provides implementations or reimplementations of 14 widely used counterfactual methods. Using this standardized setup, we conduct a comprehensive quantitative comparison across a variety of methods on datasets that differ in size, number, and types of attributes. The evaluation protocol incorporates multiple complementary metrics capturing validity, coverage, sparsity, proximity, and distributional plausibility, including density- and outlier-based measures to assess the realism of generated counterfactuals. To the best of our knowledge, this is the first comprehensive benchmark that systematically evaluates recent counterfactual explanation methods within a unified and reproducible framework. While prior libraries and benchmarking efforts exist in the literature, many are outdated, limited in scope, or lack consistent evaluation protocols. The proposed benchmark aims to improve reproducibility, enable fair comparison, and establish a workbench for the development of future counterfactual explanation methods.

Comments16 pages, 5 figures. Accepted for presentation at the XKDD and Beyond Workshop (non-archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑