审计决定裁决:工具效应在LLM决策审计中与人口统计偏见相抗衡
The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits
查看机构详情
- Carnegie Mellon University(卡内基梅隆大学)
- Amazon Web Services AI Native(亚马逊网络服务AI原生部门)
- Indian Institute of Technology(印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过大规模实验发现,LLM决策审计结果主要受审计工具设计影响,而非人口统计偏见,且此结论在招聘、贷款和医疗分诊中均成立。
中文摘要 AI 辅助
语言模型是否表现出人口统计偏见,可能取决于审计如何提出其问题。一项慈善援助基准报告称,相同的模型在逐一评分请求时偏向少数族裔申请人,而在并排排名时则对某些申请人不利。我们测试了这种反转是否适用于招聘、贷款和医疗分诊:对五个模型进行了40,726次请求,申请仅因申请人姓名不同,且主要测试在收集前已固定。结果并非如此。36个计划对比中没有一个在修正后仍然显著。评分优势保持其符号,约为已发表效应大小的一半,而精度扩展将任何招聘排名惩罚限制在已发表效应之下,尽管贷款和分诊排名下限高于该界限,因此排除结论仅对招聘排名以及所有三个领域的评分是决定性的。植入的差异跟踪其注入的大小,以及对原始援助材料的定向复制限制了这些零结果。审计比人口统计更活跃:模型几乎总能识别透明的审计,无论变化的细节是种族还是爱好,都会将每个相同内容的比较联系起来,并且对首位候选人的奖励与我们测量的任何人口统计效应一样大。审计裁决更多地反映审计构建而非人口统计偏见。
英文摘要
Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.