arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

偏见审计能检测偏见但在排名上存在分歧:来自十种工具和十个前沿模型的证据

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, José O. Gomes

arXiv 2609.15995首次发表:更新:

发表机构

Tsinghua University; Federal University of Rio de Janeiro(清华大学; 里约热内卢联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过统一推理网关对十个前沿模型运行十种偏见审计工具,发现检测成功但跨工具排名一致性等同随机,表明不同工具测量不同构念,单一审计无法支持模型间排名。

AI 中文摘要

新兴的人工智能法规要求对高风险系统进行偏见审计,审计分数开始被用于对模型进行排名。这两种用途都假设不同的审计工具能够充分测量同一事物以进行比较。我们直接检验了这一假设,通过一个统一的推理网关,对十个前沿模型的共享面板运行十种外部审计工具,首先针对职业性别偏见,然后针对年龄和社会经济地位。检测成功而排名失败。十种工具中有八种检测到偏见,置信区间不包含零;两个被广泛引用的直接探针基准已饱和,因为前沿模型现在给出中性回答。但跨工具的排名一致性无法与随机区分(Kendall's W=0.07, p=0.83)。一个包含六个故意较弱模型的阳性对照区分了两种解释:一旦面板跨越真实的能力差距,工具内的可靠性得以恢复,但跨工具排名从未恢复,这表明工具测量的是不同的构念,而非同一构念的噪声。甚至偏见的方向也因审计格式而异:强制选择决策工具大多过度纠正(偏向女性,在278个招聘决策中有273个偏向工人阶级候选人),而自由生成和默认共指仍与刻板印象一致。该模式在社会经济地位上重复;年龄上的表面排名一致性在论文自身的工具包含规则下瓦解。实际信息是:单一审计可以检测偏见并估计其在其自身操作化范围内的方向,但没有任何单一审计支持将一个模型与另一个模型进行排名。所有原始响应、代码以及从源重新计算每个报告数字的分析均可在该https URL获取。

英文摘要

Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status. Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83). A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent. The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at https://github.com/williamguey/bias-audit-agreement.

Comments18 pages, 7 figures, 4 tables. Code and data: https://github.com/williamguey/bias-audit-agreement

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑