性别偏见在大语言模型中普遍存在且高度异质
Gender bias across LLMs is common and highly heterogeneous
- University of Milan-Bicocca(米兰比可卡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过两项实验评估十个LLM的性别偏见,发现偏见普遍存在但方向和程度高度异质,强调偏见审计需持续进行并覆盖多供应商。
AI中文摘要:
随着大语言模型(LLMs)嵌入具有实际后果的决策支持工具,理解其中的性别偏见变得越来越重要。以往研究仅关注少量模型,未充分揭示性别偏见在LLMs中的普遍性和异质性。我们针对2025年4月至2026年6月间发布的、来自九家供应商的十个模型,采用两种范式来填补这一空白:对刻板短语的性别归属(研究1)以及对为防止灾难性后果而虐待或酷刑女性或男性的道德判断(研究2)。在研究1中,十个模型中有两个模型将男性刻板短语归属于女性作者的情况多于相反情况,而三个模型则表现出相反模式。在研究2中,多个模型收敛于一种不利于男性的不对称性,其方向与已有的人类倾向于保护女性目标免受伤害的倾向一致,尽管这种不对称性出现的具体条件因模型而异;相比之下,另外三个模型在不同条件下未表现出任何差异。这些结果表明,与性别相关的偏见在LLMs中普遍存在。然而,其方向和程度高度异质,以至于某些模型的行为与其他模型截然相反。因此,偏见审计应被视为一个持续的、多供应商的过程,而非一次性的评估。
英文摘要:
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.