arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30918cs.LGcs.AI

对哪种模型变化具有鲁棒性?稳健反事实解释的统一评估

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

发表机构弗罗茨瓦夫理工大学 · Tooploox
查看机构详情
  • Wrocław University of Science and Technology(弗罗茨瓦夫理工大学)
  • Tooploox

机构由 AI 辅助整理,请以论文原文为准。

Marcin Kostrzewa, Maciej Zięba

首次发表
浏览论文内容

中文总结 AI 辅助

针对反事实解释的稳健性评估缺乏统一标准,本文提出跨家族评估协议,在固定事实与反事实下测试八种模型变化,发现不同变化下性能差异显著,RobX迁移最稳定,并主张通用评估协议。

中文摘要 AI 辅助

稳健的反事实解释承诺在模型发生变化后仍能提供可行的行动方案。这些承诺是否兑现取决于变化的具体类型。参数的小幅扰动、基于新数据的重新训练以及新的架构是不同的事件,每种现有方法都是针对其设计时所针对的变化类型进行评估的。因此,报告的稳健性分数回答的是不同的问题,无法相互比较。我们提出了一种统一的跨家族评估协议,该协议固定事实实例和生成的反事实,同时针对相同的八种模型变化类型测试每种方法。该基准在四个表格数据集上比较了六种稳健方法和两种标准基线。它通过输出表征每个变化后的分类器,并报告经验稳健性以及覆盖率、基础有效性和邻近性。我们发现,相对性能和失败模式在不同变化家族之间存在差异。有界参数扰动平均改变0.95%的测试预测,而自助重训练则改变4.9%。针对这些扰动具有保证的方法不一定能迁移到其他变化。在我们的实验中,RobX的迁移最为一致,尽管更高的稳定性可能需要更大的干预。我们认为,稳健的CFE方法应通过一个通用协议进行评估,该协议明确模型变化,衡量其实现的行为幅度,并将生成性能与稳健性分开考量。

英文摘要

Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.

↑