并非对所有人安全:审计文本到图像安全流程中的方言惩罚
Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究发现文本到图像安全流程存在方言惩罚问题,五种英语方言的23080对提示实验显示,不同安全过滤器对不同方言的检测偏差显著,可通过组平衡重训练缓解。
AI中文摘要:
文本到图像(T2I)安全护栏无法公平地推广到非标准方言。我们评估了五种英语方言的23080对提示,将这种失败形式化为方言惩罚,即过滤器根据语言表面特征而非语义意图触发。文本级过滤器以相反方向失效:NSFW-T过度标记良性方言提示,LatentGuard过度标记有毒提示(偏差差距高达+28.29个百分点),而OpenAI Moderation API则检测不足。受控错别字消融实验证实,该惩罚源于对方言特征的标记,而非通用分布外敏感性。像素级生成器在很大程度上与方言无关;惩罚进入文本处理环节,并不均衡地级联到事后护栏。我们表明,这种偏差源于训练数据不平衡,可通过组平衡重训练缓解,消融实验将增益归因于平衡暴露,而非GroupDRO(组分布鲁棒优化)的最差组目标。当前流程系统性地让方言使用者失望,这一公平性失败被平均准确率基准掩盖。我们的官方代码和数据集可在此https URL公开获取。内容警告:本文包含冒犯性、有毒或令人不安的文本提示和生成图像。
英文摘要:
Text-to-image (T2I) safety guardrails fail to generalize equitably to non-standard dialects. Evaluating 23,080 paired prompts across five English dialects, we formalize this failure as the dialect penalty, where filters trigger based on linguistic surface features rather than semantic intent. Text-level filters fail in opposing directions: NSFW-T over-flags benign dialect prompts and LatentGuard over-flags toxic ones (bias gaps up to +28.29 pp), while the OpenAI Moderation API under-detects them. A controlled typo ablation confirms this penalty originates from flagging dialectal features, not generic out-of-distribution sensitivity. The pixel-level generator is largely dialect-agnostic; the penalty enters at text processing and cascades unevenly to post-hoc guardrails. We show this bias tracks training data imbalance and is mitigable via group-balanced retraining, with an ablation attributing the gain to balanced exposure rather than to the worst-group objective of GroupDRO (group distributionally robust optimization). Current pipelines systematically fail dialect speakers, an equity failure masked by mean accuracy benchmarks. Our official code and dataset are publicly available at https://github.com/minguinho26/dialect-penalty-t2i. Content Warning: This paper contains offensive, toxic, or disturbing text prompts and generated images.