arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27905cs.LGcs.AI

用于反事实解释生成的类别感知强化学习

Class-Aware Reinforcement Learning for Counterfactual Explanation Generation

发表机构卡拉奇工商管理学院 · 澳大利亚联邦银行 · 斯坦福大学
另 1 家 · 查看机构详情
  • Institute of Business Administration Karachi(卡拉奇工商管理学院)
  • Commonwealth Bank of Australia(澳大利亚联邦银行)
  • Stanford University(斯坦福大学)
  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

Muhammad Adil Saleem, Syed Ali Raza, Mary-Anne Williams

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出类别感知强化学习方法,将实例预测类别纳入RL状态表示,经7个不同领域数据集对比实验,该方法收敛更快、奖励优化更好、有效反事实解释更多,凸显类别感知的重要性。

中文摘要 AI 辅助

反事实解释(CFEs)通过生成具有调整后特征值的替代实例来实现对比结果,从而增强黑箱模型的可解释性。强化学习(RL)为CFE生成提供了一种有前景的方法,能够高效探索反事实实例,同时确保对有效性、稀疏性和 proximity( proximity 此处保留英文,指与原始实例的邻近度)等关键指标的控制。以往研究仅使用来自监督数据集中预测变量的特征来构建RL状态。本研究探讨在RL状态表示中加入实例的预测类别以及来自预测变量的特征,对生成CFE的影响,假设类别感知可提升探索效率并改善策略最优性。我们将提出的类别感知RL方法与类别盲RL方法进行对比,类别盲RL方法与前者类似,但在状态表示中排除了实例的类别信息。对比使用了来自不同领域、规模各异的7个数据集。结果显示,训练期间,类别感知RL在收敛速度、奖励优化和回合长度缩短方面具有优势;此外,与类别盲RL相比,它生成的有效CFE显著更多。最后,通过SHAP和LIME值分析表明,实例的类别相关特征始终是RL动作选择中最具影响力的预测变量之一,凸显了类别感知在用于CFE生成的RL中的重要性,其带来的效果是清晰度提升、学习速度加快、有效性改善,以及在不同数据集上生成更有效的反事实。

英文摘要

Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values that achieve a contrastive outcome. Reinforcement learning (RL) offers a promising approach for CFE generation, enabling efficient exploration of counterfactual instances while ensuring control over key metrics like validity, sparsity, and proximity. Previous studies have formulated RL states exclusively using features derived from the predictors in the supervised dataset. This study explores the impact of including an instance's predicted class, alongside features derived from the predictors, in the RL state representation for generating CFEs. The hypothesis is that class-awareness enhances exploration efficiency and improves policy optimality. We compare the proposed class-aware RL method with the class-blind RL method, which is similar but excludes the instance's class information from the state representation. The comparison was conducted using seven datasets from diverse domains, varying in size. The results show that during training, class-aware RL offers benefits in terms of convergence speed, reward optimization, and episode length reduction. Moreover, it generates significantly more valid CFEs compared to class-blind RL. Finally, the instance's class-based feature consistently ranks among the most influential predictors in RL's action-selection, as shown by the SHAP and LIME values, underscoring the significance of class-awareness in RL for CFE generation. The impact is heightened clarity, faster learning, improved validity, and more effective counterfactual generation across diverse datasets.

↑