AI 中文总结
对巴塞罗那公共就业机构Barcelona Activa的TalentClue半自动化招聘系统开展端到端公平性审计,发现其存在性别、年龄、非二元性别等维度的公平性差异,凸显需开展社会技术层面的端到端评估。
AI 中文摘要
算法公平性评估通常将AI系统作为有界技术组件进行评估,忽略了其运行所处的组织背景。我们开展了据我们所知,对公共就业机构Barcelona Activa运营的半自动化招聘系统的首次独立端到端公平性审计,该机构使用第三方TalentClue平台进行候选人搜索和初选。我们分析了2017年9月至2022年9月期间约497000条候选人-职位管道条目,涵盖了自动处理、人工决策、候选人数据和雇主决策在内的七个管道阶段。按二元性别划分的总体结果在统计上无差异,但这种均等掩盖了薪资水平、年龄和性别认同方面的巨大差异。女性在中等薪资初选阶段面临不利影响(差异比率DIR=0.786,p<0.001),20个行业中有15个存在薪资差异,46-55岁女性还面临叠加劣势(DIR=0.77)。非二元性别候选人的初选率不到男性的三分之一(DIR=0.295),不过该估计基于小样本(N=285)。55岁及以上候选人完全未出现在管道中,尽管这一群体占巴塞罗那劳动力的15.6%。初选的性别差距随时间缩小,从2017年的6.5个百分点降至2022年的1.3个百分点。审计还揭示了供应商与部署者之间的信息不对称:Barcelona Activa无法获取TalentClue匹配逻辑和评估的关键信息。公平性结果可产生于自动处理、人工决策、数据质量、供应商透明度和管道结构之间的相互作用。我们基于此前对社会技术、端到端公平性评估的呼吁,实证表明仅靠模型层面评估不足以理解已部署系统的公平性。
英文摘要
Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 to September 2022, covering seven pipeline stages that span automated processing, human discretion, candidate data, and employer decisions. Aggregate outcomes across binary genders are statistically indistinguishable, yet this parity masks substantial disparities by salary level, age, and gender identity. Women face adverse impact in mid-salary shortlisting (DIR = 0.786, p < 0.001), alongside salary disparities in 15 of 20 sectors and a compounded disadvantage for women aged 46-55 (DIR = 0.77). Non-binary candidates are shortlisted at less than one third the rate of men (DIR = 0.295), although this estimate rests on a small sample (N = 285). Candidates aged 55 and over are entirely absent from the pipeline despite comprising 15.6% of Barcelona's labor force. The gender gap in shortlisting narrows over time, from 6.5 percentage points in 2017 to 1.3 in 2022. The audit further reveals a vendor-deployer information asymmetry: Barcelona Activa lacks access to key information about TalentClue's matching logic and evaluation. Fairness outcomes can thus arise from interactions among automated processing, human discretion, data quality, vendor opacity, and pipeline structure. We build on prior calls for sociotechnical, end-to-end fairness evaluation, showing empirically why model-level assessment alone can be insufficient for understanding fairness in deployed systems.