arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26955cs.LGcs.CY

后处理公平性约束何时有帮助、何时有害:来自八项跨领域评估的证据

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

Nithin Raghava Ramachandra Narla

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出FAPE框架,通过八项跨领域评估发现后处理公平性约束在高差异场景下有效但可能恶化接近公平场景,主张采用基线差异筛查与持续监控替代单一部署时审计。

中文摘要 AI 辅助

生产环境中的机器学习公平性审计通常只在部署时进行一次,且仅针对单一领域。这两种做法在实践中都会失败:公平性可能在重新训练或用户群体变化后发生偏移,而在一个数据集上验证的干预措施很少会在组织部署的异构领域中进行测试。我们提出了FAPE(生产环境公平性审计框架),这是一个四阶段框架,用于评估单一的後处理干预措施——Fairlearn的ThresholdOptimizer,涵盖八项领域评估:刑事司法、收入预测、法律录取、信贷借贷、农业借贷、多领域基准语料库、医疗保健和教育。每项评估均以人口统计均等性和均等化几率差异进行评分,并在可计算的情况下辅以差异影响比和准确率成本。干预效果与基线差异幅度相关:在模型-领域配对中,约束在14个高差异案例中的9个改善了差异,而在4个接近公平的案例中的3个使其恶化。五个高差异例外中的每一个都在两种测量检查之一(最小群体规模或在保留数据上拟合的阈值)下发生逆转。一个从部署时启动的CUSUM监控器,在模拟偏移上进行测试,将从未达到0.1均等约定的受约束模型与达到该约定但后来发生回归的模型区分开来。因此,单一的部署时审计是不可靠的指南,这主张进行基线差异筛查和持续监控。

英文摘要

Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We present FAPE (Fairness Auditing for Production Environments), a four-stage framework evaluating a single post-processing intervention, Fairlearn's ThresholdOptimizer, across eight domain evaluations: criminal justice, income prediction, legal admissions, credit lending, agricultural lending, a multi-domain benchmark corpus, healthcare, and education. Each is scored on demographic parity and equalized odds difference, plus disparate impact ratio and accuracy cost where computable. Intervention effectiveness tracks baseline disparity magnitude: across model-domain pairs the constraint improved disparity in 9 of 14 high-disparity cases and worsened it in 3 of 4 near-fair ones. Each of the five high-disparity exceptions reverses under one of two measurement checks, a minimum group size or thresholds fit on held-out data. A CUSUM monitor started at deployment, tested on a simulated shift, separates constrained models that never met a 0.1 parity convention from those that met it and later regressed. A single deployment-time audit is therefore an unreliable guide, which argues for baseline-disparity screening and continuous monitoring

补充信息

↑