面向自动招聘的反事实逐决策偏差审计:定位与解释申请人跟踪系统中的差异影响
Counterfactual, Per-Decision Bias Auditing for Automated Hiring: Localizing and Explaining Disparate Impact in Applicant Tracking Systems
浏览论文内容
中文总结 AI 辅助
本文提出AI偏差防火墙(AIBF),用于逐决策审计自动招聘的申请人跟踪系统,在真实数据集上验证其能精准定位偏差决策与受损害候选人,但纠正偏差无法达到法律要求的公平性。
中文摘要 AI 辅助
申请人跟踪系统(ATS)越来越多地决定招聘过程中哪些候选人进入下一环节,当前诉讼与监管要求这些决策具备可审计性。现有工具分为两类:差异影响比率等群体公平指标可汇总整个人群的情况,但无法说明哪些个体决策不公平或原因;SHAP等局部解释器可归因于单个预测,但与招聘偏差判定的法律标准不相关。本文提出AI偏差防火墙(AIBF),这是一种逐决策审计申请人跟踪系统的方法:AIBF会中和候选人受保护属性的代理变量,重新评分决策,并测量由此产生的反事实偏移,从而得出以分数点为单位的带符号逐决策偏差、受保护属性改变决策的标记,以及说明责任因素的通俗语言解释。我们在两个真实公开数据集Adult和COMPAS上进行评估,而非合成数据。逐决策反事实偏移具有一致性,汇总后可重现已知的群体层面差异,例如在Adult数据集上,特权群体的平均偏移为+7.5分,弱势群体为-8.0分,与测得的统计 parity差异一致。AIBF以0.963的ROC曲线下面积识别受保护属性翻转的决策,而按群体成员标记的基线仅为0.672;它识别受损害候选人的精度极高,仅审查5%的决策就能发现55%的受损害候选人,而基于群体审查的比例仅为6%。我们还报告了一个局限性:纠正标记的决策会大幅提高差异影响比率,但无法达到法律层面的 parity,因为标记为绩效的特征仍存在残余代理相关性。AIBF以Apache 2.0许可证发布,附带代码与实验。
英文摘要
Automated applicant tracking systems increasingly decide who advances in hiring, and litigation and regulation now demand that those decisions be auditable. Existing tools sit at two extremes. Group fairness metrics such as the disparate impact ratio summarize a whole population but cannot say which individual decisions were unfair or why, while local explainers such as SHAP attribute a single prediction but are not connected to the legal standard by which hiring bias is judged. We present the AI Bias Firewall (AIBF), a method that audits an applicant tracking system one decision at a time. AIBF neutralizes a candidate's protected-attribute proxies, re-scores the decision, and measures the resulting counterfactual shift, which yields a signed per-decision bias in score points, a flag for decisions the protected attributes changed, and a plain-language explanation naming the responsible factors. We evaluate on two real public datasets, Adult and COMPAS, rather than on synthetic data. The per-decision counterfactual shift is faithful, aggregating to reproduce the known group level disparity, for example a mean shift of +7.5 points for the privileged group and -8.0 for the disadvantaged group on Adult, consistent with the measured statistical parity difference. AIBF identifies the decisions that protected attributes flipped with an area under the ROC curve of 0.963 on Adult, against 0.672 for a baseline that flags by group membership, and it identifies the harmed candidates so precisely that reviewing only five percent of decisions surfaces fifty-five percent of them, against six percent under group based review. We also report a limitation: correcting flagged decisions raises the disparate impact ratio substantially but not to legal parity, because features labeled as merit carry residual proxy correlation. AIBF is released under the Apache 2.0 license with code and experiments.