反事实解释与可争议性的范围
Counterfactual Explanations and the Scope of Contestability
浏览论文内容
中文总结 AI 辅助
本文探讨反事实解释是否有助于质疑算法决策,构建可争议性的阐释,分析其作用局限,并提出可使其更适配的多轮反事实方法,以恢复决策主体的自主性。
中文摘要 AI 辅助
在社会领域中,不透明的机器学习模型对具有重要后果的决策实现自动化,这阻碍了我们的自主性。本文探讨如何通过提供特定类型的知识来恢复这种自主性,更准确地说,我们讨论一种特定类型的解释——反事实解释,是否有助于我们质疑算法决策。在此背景下,本文作出三项贡献:第一,我们构建了可争议性的阐释,将其定义为向决策主体提供足够信息,使其能以此为基础要求撤销决策,同时我们将可争议性与解释权相关论述中的邻近概念(如正当性和追索权)划清界限。第二,我们通过考量导致有问题的算法决策的多种失败模式,审查反事实解释在多大程度上有助于实现可争议性,并仔细审视其在多大程度上帮助我们检测潜在错误。第三,我们提出通过某些修改,可使反事实解释更适合服务于预期功能,在此过程中,我们勾勒出一种多轮反事实方法的轮廓,决策主体可在(有限)次数内查询模型以测试自身的反事实假设。
英文摘要
The automation of consequential decisions through opaque machine learning models in societal domains impedes our agency. This paper is about how agency can be reinstated by the provision of certain kinds of knowledge. More precisely, we discuss whether a specific type of explanation, counterfactual explanations, facilitates our ability to contest algorithmic decisions. Against this backdrop, our paper makes three contributions: First, we develop an account of contestability, where contestability is defined as the provision of information, sufficient for a decision-subject to use as a basis for demanding that a decision be revoked. We also demarcate contestability from adjacent concepts in the discourse surrounding the right to explanation, such as justification and recourse. Second, we examine to what extent counterfactual explanations are conducive to contestability by considering a variety of failure modes causing problematic algorithmic decisions and scrutinize to what extent counterfactual explanations help us detect the underlying errors. Third, we propose ways in which, with certain modifications, counterfactual explanations can be made more fitting to serve the desired function. In this vein, we sketch the contours of a multi-shot approach to counterfactuals, where decision-subjects can query a model to test their own counterfactuals for a (limited) number of times.