arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18106cs.CLcs.AIcs.CY

开源权重大语言模型在招聘应用中引发性别与种族偏见的语言触发因素

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

  • Recruit Co., Ltd.(瑞可利有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Kosuke Kitahara, Nobuhiro Yamaguchi

AI总结:

本研究首次系统审计六个开源权重LLM,发现招聘语言中的代理型与编码排斥词汇分别抑制女性和非白人候选人评分,并提出部署前审计协议以应对欧盟AI法案合规风险。

AI中文摘要:

开源权重的大语言模型正迅速进入招聘流程,但其歧视性故障模式——以及这些模式在欧盟《人工智能法案》高风险分类(附件三)和美国平等就业机会委员会不利影响分析下所产生的监管风险——仍鲜为人知。我们首次对开源权重大语言模型进行了系统性、多模型的审计,将招聘岗位语言作为主要实验变量,评估了六个模型(Llama 3.2、Mistral、Gemma 3、Qwen 3、Phi 3、DeepSeek-R1),通过四项受控实验共同探究招聘者模拟和求职者模拟任务。我们发现:(1)代理型岗位语言降低了女性候选人的招聘者推荐分数(r_rb = 0.309,p_Bonf = 7x10^-5;模型固定效应 r_rb = 0.448),而社群型语言部分逆转了这一惩罚;(2)编码排斥语言在较大效应量下(r_rb = 0.646-0.758)抑制了非白人招聘者的评分,并且在求职者方面,选择性地阻止非白人角色表达兴趣——在大规模上实现了寒蝉效应机制。标签消融实验将显式的人口统计角色标签隔离为主要因果驱动因素,词嵌入关联测试在表征层面佐证了这些发现(在Caliskan等人的多词性别属性列表下,d = 1.01-1.45)。我们将这些结果转化为具体的部署前审计协议——岗位词汇评分、角色条件化LLM探测以及针对五分之四阈值的不利影响标记——该协议落实了附件三对招聘中高风险人工智能所施加的记录和风险管理义务。

英文摘要:

Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.

补充信息

↑