大型语言模型中作为地理偏见的机构声望:来自三项带自助法置信区间的析因实验的证据
Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
浏览论文内容
中文总结 AI 辅助
该研究通过三项带自助法置信区间的析因实验,发现大型语言模型在候选人评估中存在机构声望、期刊声望及国籍相关的偏见,期刊声望影响远大于机构声望,还证实了发表于《自然》的救助效应。
中文摘要 AI 辅助
我们研究大型语言模型(LLMs)在候选人评估中是否会基于申请者姓名的族裔、机构声望及地理位置进行系统性歧视。报告了三项析因实验(共4320次API调用、涉及4个LLMs、5个专业领域)。研究1(3×4设计)发现,在10分制下存在统计学上稳健的机构层级梯度,为+0.297分(95%自助法置信区间:+0.175至+0.422),而姓名起源效应可忽略且无统计学意义(95%置信区间包含零)。研究2(2×2声望×国家设计)拆分了声望与地理的混淆因素:声望效应(+0.185;95%置信区间:+0.093至+0.275)比国籍效应(+0.126;95%置信区间:+0.037至+0.218)高出1.5倍。研究3(2×2期刊×机构设计)显示,期刊声望(《自然》与某边缘开放获取期刊)对机构声望的影响占比达5.7倍:期刊效应为+1.937(95%置信区间:+1.811至+2.062),而机构效应为+0.341(95%置信区间:+0.184至+0.504)。还证实了“救助效应”:来自瓜亚基尔大学的申请者,发表于《自然》期刊对低机构声望的补偿作用更强(+2.127),而麻省理工学院申请者的该补偿值为+1.745。结果使用中智偏见指数NBI<T,I,F>量化;其中I分量显示,低声望背景的候选人评估不一致性更高,这是仅用均值指标无法捕捉的认知劣势。代码与数据:此httpsURL
英文摘要
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI<T,I,F>; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm
发表机构
- Universidad Bolivariana del Ecuador(厄瓜多尔玻利瓦尔大学)
- Universidad de Guayaquil(瓜亚基尔大学)
- University of New Mexico(新墨西哥大学)
机构由 AI 辅助整理,请以论文原文为准。