发表机构
Durham University(杜伦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示 GPT 系列模型安全训练中,显性性别歧视被转化为隐性表征危害而非消除,提出“危害洗白”概念及三标准检测协议,证明毒性分数下降不足以代表危害减少。
AI 中文摘要
大型语言模型的安全评估依赖于表面形式的分类器,这些分类器报告跨模型代际的危害分数下降。我们提供的证据表明,这种方法论存在系统性缺陷:显性歧视内容被转化而非被移除。我们称之为“危害洗白”。通过分析跨越 GPT-2 至 GPT-5(OpenAI GPT 系列;三种人口统计条件)的 15 个模型中的 450,000 个面向性别的补全,我们表明在 GPT-2 中面向女性的输出中普遍存在的性暴力聚类到 GPT-4 时消失,而面向男性的补全获得了正向表征领域(如照护、情感范围、盟友身份),这些在面向女性的输出中并不存在。该模式在 GPT-5 中最为明显:主题 5(1,997 个文档)将乳腺癌框定为男性权利辩论,而面向女性的输出中未出现等价聚类。三个独立分类器将此内容评为无毒。情感分数在 GPT-4 处反转:早期模型贬低女性;后期模型过度纠正。在 GPT-4 对齐边界处,面向女性的补全的主题多样性相对于男性下降 36%(W/M = 0.58,而 GPT-2 时为 0.91)。REGARD 表征危害差异与发布日期相关(ρ = +0.55,p = .034),而 Detoxify 不相关(ρ = -0.23,p = .42):毒性分数下降而表征危害增长。我们将危害洗白形式化为三标准测试,并提供适用于任何生成模型的三阶段检测协议。在 OpenAI GPT 系列中,毒性分数降低不足以作为危害减少的充分代理。
英文摘要
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($ρ= +0.55$, $p = .034$) while Detoxify does not ($ρ= -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.
CommentsAccepted at EMNLP 26 Main Conference