arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16501cs.CLcs.AI

超越名称:去标识化简历中的人口统计泄露与LLM偏见审计中的评估伪影

Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits

  • Macquarie University(麦考瑞大学)
  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

Qiangju Chen, Yang Xiao

AI总结:

本研究通过控制语言变量,证明非语言文本仍可泄露人口统计信息,且评估设计显著影响LLM偏见审计结果,凸显了区分内容信号与协议效应的必要性。

AI中文摘要:

去标识化的简历筛选假设删除显式字段可以防止族裔文化推断;然而,最近的审计将残留泄露归因于声明的语言。我们研究了消除语言字段是否能解决这一泄露问题,涉及九个开放权重模型和620份反事实简历。通过严格保持语言属性一致,我们隔离了五种族裔文化条件和三种线索显著性层级下的非结构化文本。目标群体恢复平均为0.757,在高显著性下饱和至1.000,表明非语言文本足以维持人口统计推断。关键的是,模型仅在微弱线索下出现分歧(0.086-0.690),确立了显著性作为评估的必要维度。此外,成对的LLM作为评判者的结果对评估设计高度敏感:禁止平局产生明显的选择率比值为0.39,同时伴随强烈的位置和内容效应,而允许平局则导致大多数模型近乎普遍的平局(≥94%)。下游评分仅显示条件间非常小的差异,强调需要区分可从简历内容中恢复的人口统计信号与评估协议引入的效应。

英文摘要:

De-identified résumé screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual résumés. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from résumé content from effects introduced by the evaluation protocol.

补充信息

↑