发表机构
Beijing Normal University(北京师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对不同审计目标,提出统一的基于容限的公平性审计框架,通过约束经验似然检验和分裂经验似然检验实现违规认证与敏感性筛查,并在实验中展示了误差控制与敏感性的权衡。
AI 中文摘要
随着人工智能的日益部署,算法不公平性引发了越来越多的关注,并加强了对透明公平性审计的需求。在实践中,算法不公平的可容忍程度取决于具体的法律、伦理或应用背景。给定一个预先指定的容限阈值,一个重要的统计问题是,在不同的审计目标下,如何确定群体差异是否超过允许的容限。为解决这一问题,我们针对两个互补的审计目标开发了一个统一的基于容限的公平性审计框架:违规认证,其优先控制虚假违规声明的比率;以及敏感性筛查,其优先减少遗漏的违规。对于第一个目标,我们开发了一种用于正式场景的约束经验似然检验,该检验采用最不利点校准,并可结合虚假标记率控制以进行同时的子群体审计。对于第二个目标,我们开发了分裂经验似然和调整后的分裂经验似然检验,采用自适应边界代理原则用于预警场景。数值实验展示了这些程序在错误控制与敏感性之间的不同权衡。一项COMPAS分析说明了该框架在预测性公平性审计中的应用。
英文摘要
As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairness auditing. In practice, the tolerable degree of algorithmic unfairness depends on the specific legal, ethical, or application context. Given a prespecified tolerance threshold, an important statistical question is how to determine whether a group disparity exceeds the allowable tolerance across different auditing objectives. To address this problem, we develop a unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations. For the first objective, we develop a constrained empirical likelihood test for formal settings that uses least-favorable-point calibration and can be combined with false flagging rate control for simultaneous subgroup auditing. For the second objective, we develop split empirical likelihood and adjusted split empirical likelihood tests using an adaptive boundary-proxy principle for early-warning settings. Numerical experiments show the distinct error-control--sensitivity trade-offs of these procedures. A COMPAS analysis illustrates the framework in predictive fairness auditing.
Comments47 pages, 8 figures