什么使公平差距具有可操作性?面向负责任AI部署的统计可操作性
What Makes a Fairness Gap Actionable? Statistical Actionability for Responsible AI Deployment
浏览论文内容
中文总结 AI 辅助
该研究提出统计可操作性框架,整合多维度公平性证据生成部署建议,在模拟中降低决策成本,可区分相似公平差距审计,为负责任AI部署提供决策层支持。
中文摘要 AI 辅助
算法公平性审计能够检测出差异,但无法确定这些差异何时值得干预。部署决策还取决于证据的可靠性、子群体支持情况以及部署环境。现有公平性方法可量化差异与不确定性,但在将积累的证据转化为行动方面提供的指导有限。我们提出统计可操作性(Statistical Actionability),这一统计概念将公平性部署重新定义为基于证据的决策问题。该框架整合了关于差异幅度、统计可靠性、子群体充足性及部署环境的公平性证据,并将所得证据状态映射至四类建议:缓解、收集更多数据、监测或不立即采取行动。在受控模拟中,统计可操作性在代表性基线中实现了最低决策成本,相较于基于差距的干预方式,平均决策成本降低19.2%,同时减少了误报和漏报偏差。校准后的部署规则在异质统计环境中具有泛化能力,在五种可迁移性场景中有四种场景下与目标神谕的偏差保持在2%以内。对基准公平性审计的分析进一步表明,该框架可区分具有相似观测公平差距但不确定性和子群体支持水平不同的审计,生成可解释的部署建议。因此,统计可操作性在公平性评估与部署干预之间建立了一个统计决策层,使负责任的AI系统能够依据积累的证据而非仅差异幅度采取行动。
英文摘要
Algorithmic fairness audits can detect disparities, but they do not determine when those disparities warrant intervention. Deployment decisions also depend on the reliability of the evidence, subgroup support, and deployment context. Existing fairness methods quantify disparities and uncertainty, yet provide limited guidance for translating accumulated evidence into action. We introduce Statistical Actionability, a statistical construct that recasts fairness deployment as an evidence-based decision problem. The framework integrates fairness evidence regarding disparity magnitude, statistical reliability, subgroup adequacy, and deployment context, and maps the resulting evidence state to one of four recommendations: mitigate, collect more data, monitor, or take no immediate action. In controlled simulations, Statistical Actionability achieved the lowest decision cost among representative baselines, reducing average decision cost by 19.2% relative to gap-based intervention while simultaneously reducing both false alarms and missed bias. A calibrated deployment rule generalized across heterogeneous statistical environments, remaining within 2% of the target oracle in four of five transportability regimes. Analyses of benchmark fairness audits further demonstrated that the framework distinguished audits with similar observed fairness gaps but different levels of uncertainty and subgroup support, yielding interpretable deployment recommendations. Statistical Actionability therefore establishes a statistical decision layer between fairness evaluation and deployment intervention, enabling responsible AI systems to act on accumulated evidence rather than disparity magnitude alone.