发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FairFund-Bench基准通过调整审计格式等特征,发现LLM资源分配的偏差受审计类型影响,且因果框架效应强于人口统计效应,可用于评估LLM的分配偏差。
AI 中文摘要
大语言模型(LLM)越来越多地参与稀缺资源的分配,引发了对基于种族、性别等特征产生偏差分配的担忧。然而,近期对LLM的审计结果不一致,即便针对同一模型,也发现了对女性和少数族裔存在正向与负向歧视的证据。我们表明这种分歧可能源于审计格式的差异,并引入FairFund-Bench这一基准,该基准系统地改变了过往审计设计的关键特征:评估任务(评分、排序或分配)、比较语境(单刺激或多刺激),以及审计是透明还是隐蔽。该基准包含600份财务援助申请,由人工撰写的模板生成(已针对130万份真实GoFundMe筹款活动进行校准),覆盖三个领域、四个种族类别、两个性别类别,以及源自福利应得性理论的五种需求因果框架。在14个模型中,审计格式会改变偏差方向:模型在单独评分申请人时偏向少数族裔,但在并排排序时会惩罚部分群体。偏差幅度总体较小,但隐蔽审计中的偏差是透明审计的数倍,在透明审计中,面对仅申请人姓名不同的申请,模型会压倒性地平均分配资金。相比之下,因果框架效应比人口统计效应高出约一个数量级,且在所有模型和审计格式中一致,表明当前LLM会稳健地重现人类的应得性评估。该基准从四个标准(人口统计偏差、应得性对齐、跨任务一致性、跨语境一致性)对模型评分,公开可用,且可轻松适配其他实质性领域。
英文摘要
Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-language requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 19 pages, 7 figures. Code and data: https://github.com/martinlukk/fairfund-bench