AI 中文总结
本研究通过对比GPT-5.2与Gemini 2.5 Flash在英语和斯瓦希里语上的9类人口统计偏见表现,发现偏见会转变而非转移,仅英语偏见审计无法覆盖多语言部署需求。
AI 中文摘要
大语言模型正越来越多地应用于多语言场景中,但安全对齐与偏见评估仍以英语为中心。本研究通过向GPT-5.2和Gemini 2.5 Flash提交4900组对称的英语-斯瓦希里语提示对,覆盖9个人口统计偏见维度,生成19600条补全结果,从刻板印象流行度、情感倾向、弃权(不执行)行为及跨语言语义相似度四个维度进行评估。研究发现偏见会发生转变而非转移:特定维度上刻板印象比例变化高达12个百分点,Gemini的中性情感比例在斯瓦希里语中翻倍,GPT-5.2在英语中对169条提示弃权(不执行),在斯瓦希里语中则为零,表明弃权行为在行为层面锚定英语表面形式;超过55%的提示对在两种模型中产生语义不同的补全结果。这些结果表明仅针对英语的偏见审计无法为多语言部署提供充分覆盖。
英文摘要
Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.