求职申请跟踪系统的反事实偏见测试
Counterfactual Bias Testing for Application Tracking System
浏览论文内容
中文总结 AI 辅助
该研究提出一种基于LLM智能体的自动化多指标反事实偏见审计方法,用于检测求职跟踪系统的人口统计偏见,其结果表明该方法可发现单一指标遗漏的临界公平性问题。
中文摘要 AI 辅助
自动化求职者-岗位匹配系统在新兴监管下日益被归类为高风险人工智能,但对其人口统计偏见进行审计成本高昂:经典的通信审计研究需要手动制作简历并人工提交,无法适配快速的流水线重新训练周期。本文提出一种通用、可复用的方法:(1)使用任务专用的大语言模型(LLM)智能体合成无身份偏差的基础简历,并在五个受保护特征维度(性别、年龄、居住地、语言、残疾状况)注入受控人口统计处理,生成K×(1+N)的通信审计矩阵;(2)根据符合欧盟人工智能法案的提示,定性标记推断出的受保护特征;(3)通过微调后的句子嵌入模型和余弦相似度,将求职者与岗位描述进行排名;(4)计算包含反事实(分数差值、平均绝对排名变化、翻转率)、群体公平性(前K留存率、五分之四/影响比率)以及与绩效相关(召回率@K、归一化折损累积增益@K、平等机会、平等化赔率)三类共九项指标的公平性套件,每项指标均包含自助法置信区间、显著性检验和Benjamini-Hochberg校正,最终生成包含综合风险评分的自动化PASS/INVESTIGATE/FAIL报告。在包含5个岗位、100名基础求职者和10种人口统计处理的示例语料库(共90项指标×变体评估)中,所有处理的分数变化、前K留存率和与绩效相关的比率差距均保持在可容忍范围内,但排名稳定性指标(平均绝对排名变化)和归一化折损累积增益@K各呈现出临界结果——包括中性基线本身的一项结果,这是仅查看分数或留存率的视角会遗漏的。研究结果表明,相较于任何单一综合分数,多指标、多类别审计更具优势,且LLM智能体生成的审计可作为人工策划审计的实用、低成本补充,适用于任何求职者-岗位匹配流水线。
英文摘要
Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demographic bias is expensive: classical correspondence-audit studies require hand-crafted resumes and manual submission, which does not scale to fast pipeline retraining cycles. This paper presents a general, reusable methodology that (1) uses task-specialized LLM agents to synthesize identity-neutral base resumes and inject controlled demographic treatments across five protected-characteristic axes (sex/gender, age, residence, language, disability), producing a K x (1+N) correspondence-audit matrix; (2) qualitatively flags inferred protected characteristics per an EU AI Act-aligned prompt; (3) ranks candidates against a job description via a fine-tuned sentence-embedding model and cosine similarity; and (4) computes a nine-metric fairness suite spanning counterfactual (score delta, mean absolute rank change, flip rate), group-fairness (top-K retention, four-fifths/impact ratio), and merit-aware (Recall@K, nDCG@K, equal opportunity, equalized odds) families, each with bootstrap confidence intervals, significance tests, and Benjamini-Hochberg correction, culminating in an automated PASS/INVESTIGATE/FAIL report with a composite risk score. On an example corpus of 5 job orders, 100 base candidates, and 10 demographic treatments (90 metric x variant evaluations): score shifts, top-K retention, and merit-aware rate gaps stay within tolerance for every treatment, but a rank-stability metric (MARC) and nDCG@K each surface borderline findings - including one on the neutral baseline itself - that a score- or retention-only view would miss. The results argue for multi-metric, multi-family auditing over any single aggregate score, and for LLM-agent-generated audits as a practical, low-cost complement to human-curated audits for any candidate-job matching pipeline.
发表机构
- ManpowerGroup Services India Pvt. Ltd.(万宝盛华印度服务私人有限公司)
- Indian Institute of Technology, Hyderabad(印度理工学院海得拉巴分校)
机构由 AI 辅助整理,请以论文原文为准。