发表机构
Faculty of Computer Science, Dalhousie University(达尔豪斯大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过匹配伪装范式探测十个大语言模型,发现其推理分布中沿种族与声望维度存在隐性方言偏见,且各变体受不同刻板印象簇影响。
AI 中文摘要
大语言模型(LLMs)越来越多地被部署在住房筛选等高 stakes 领域。虽然对齐技术减轻了生成文本中的显性种族偏见,但它们往往使内部概率分布中的隐性态度关联保持不变。借鉴匹配伪装社会语言学范式,我们考察了四种英语变体在住房相关社会判断中的隐性方言偏见:标准美国英语(SAE)、非裔美国白话英语(AAVE)、尼日利亚标准英语(NSE)和尼日利亚皮钦语(NP)。AAVE 反映了先前隐性偏见评估中研究的种族化方言,而 NSE 和 NP 则代表了该文献中缺失的非洲黑人后殖民变体。使用 260 个意义匹配的句子四元组和对住房相关形容词的对数概率评分,我们在三个社会亲近度不同的情境中探测了十个开放权重 LLM:租户筛选、邻居接纳和室友选择。在所有十个模型中,AAVE 和 NP 始终与比 SAE 更负面的形容词相关联,其中 NP 受到的惩罚最为严重。关键的是,每种方言都通过不同的刻板印象簇而非通用的非标准类别受到惩罚。NSE 具有制度性声望,表现出依赖情境的转变:在正式租户筛选中优于 SAE,但随着社会亲近度的增加而受到越来越大的惩罚。我们的研究结果表明,LLM 沿种族身份和声望两个维度继承了隐性方言偏见,呼应了已记录的人类住房歧视,并展示了这种偏见在后殖民英语变体中的影响范围。
英文摘要
Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions untouched. Adapting the matched-guise sociolinguistic paradigm, we examine covert dialect bias in housing-related social judgments across four varieties: Standard American English (SAE), African American Vernacular English (AAVE), Nigerian Standard English (NSE), and Nigerian Pidgin (NP). AAVE reflects the racialized dialect studied in prior covert-bias evaluations, whereas NSE and NP represent Black African, postcolonial varieties absent from this literature. Using 260 meaning-matched sentence quadruples and log-probability scoring over housing-relevant adjectives, we probe ten open-weight LLMs across three contexts varying in social proximity: tenant screening, neighbor acceptance, and roommate selection. Across all ten models, AAVE and NP are consistently associated with more negative adjectives than SAE, with NP penalized most severely. Crucially, each dialect is penalized via distinct stereotype clusters rather than a generic non-standard category. NSE, which carries institutional prestige, displays a context-dependent shift: favored over SAE in formal tenant screening but increasingly penalized as social proximity grows. Our findings reveal that LLMs inherit covert dialect bias along both racial identity and prestige dimensions, echoing documented human housing discrimination and demonstrating its reach across postcolonial English varieties.
Comments12 pages, 5 figures. Accepted to the 9th AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)