发表机构
Project AWARE(AWARE项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文以智利姓氏为探针,发现模型编码的潜在社会关联与重要决策效应解离,关联强度无法可靠预测决策泄露,表明需直接测量关联到行动的过渡。
AI 中文摘要
偏差评估常过快从模型编码社会关联的证据,推断该关联会改变重要决策。本文以智利姓氏作为受控社会经济探针,检验该推断是否成立。评估8个冻结模型提供商单元,每个单元对应1032个提示,共得到8256条验证的主要响应。设计将强制潜在关联与匹配的重要决策分离,涉及学术选拔、职业招聘、研究奖学金选拔及法律援助申请。8个模型中有7个,精英编码姓氏的强制高地位概率质量高于普通姓氏;所有8个模型中,精英编码姓氏的概率质量高于稀有频率对照。然而,多数系统中精英与普通姓氏的决策效应差值接近零。5个模型在预先声明的±0.10标准差范围内统计等价,其余3个精度不足或处于临界状态,无一致的精英优势。关联强度无法可靠预测跨模型的决策泄露(相关系数r=0.201,p=0.633),也无法可靠预测跨冻结姓氏对-模型单元的决策泄露(r=0.065,p=0.565)。核心结果是测量解离:潜在社会关联与重要待遇是经验上不同的构念,评估应直接测量从关联到行动的过渡。
英文摘要
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.