arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更新并不更公平:跨模型世代的文生图AI中的性别刻板印象

Newer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations

Shesh Narayan Gupta, Nik Bear Brown

arXiv 2609.18007首次发表:更新:

发表机构

College of Engineering, Northeastern University(东北大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究跨四代Stable Diffusion文生图模型在20个职业中的性别表现,发现76.4%图像为男性,女性占比被低估20-46个百分点,且偏差随模型更新未稳定改善,最新模型仍未实现性别均等。

AI 中文摘要

文生图生成模型在专业和创意环境中被广泛使用,然而,它们如何跨职业表现性别——以及更新模型是否更公平——在多个世代中仍知之甚少。我们评估了20个职业、5个提示模板和4个Stable Diffusion模型世代(SD 1.5、SD 2.1、SDXL、SD 3 Medium)中的性别表现,生成了8,000张图像,每个职业-模型单元格n=100(5个提示×20张图像),并使用DeepFace对所有图像进行分类。在这8,000张开源图像中,76.4%显示男性主体(95%置信区间[75.1%, 78.7%],p < 2.2×10^-16,Benjamini-Hochberg校正)。更引人注目的是,历史上女性编码职业的图像中,57.6%显示男性主体(原始p = 3.43×10^-22,BH校正p = 1.71×10^-21)。本文报告的九个显著检验在10个检验中均通过BH校正。与美国劳工统计局劳动力数据相比,模型平均低估女性20-46个百分点,对于接近性别平衡的职业偏差尤其大:科学家(BLS中女性占48%,模型输出中男性占82-99%)和清洁工(BLS中女性占46%,模型输出中男性占80-92%)。模型世代并未稳步改进:偏差从SD 1.5到SDXL恶化,然后在SD 3 Medium中部分恢复。与GPT-image-1在五个职业上的初步比较表明,其偏差低于开源模型,尽管实际效果较小(Cramer's V = 0.080),且该比较是探索性的。没有模型实现性别均等。

英文摘要

Text-to-image generative models are widely used in professional and creative settings, yet how they represent gender across occupations -- and whether newer models are fairer -- remains poorly understood across multiple generations. We evaluate gender representation across 20 occupations, 5 prompt templates, and 4 Stable Diffusion model generations (SD 1.5, SD 2.1, SDXL, SD 3 Medium), generating 8,000 images with n = 100 per occupation-model cell (5 prompts x 20 images), and classifying all with DeepFace. Across the 8,000 open-source images, 76.4% show male subjects (95% CI [75.1%, 78.7%], p < 2.2 x 10^-16, Benjamini-Hochberg adjusted). More strikingly, 57.6% of images for historically female-coded occupations show male subjects (raw p = 3.43 x 10^-22, BH-adjusted p = 1.71 x 10^-21). All nine significant tests reported in this paper survive BH correction across 10 tests. When compared against U.S. Bureau of Labor Statistics workforce data, models underrepresent women by 20-46pp on average, with particularly large deviations for near gender-balanced occupations: scientist (48% female in BLS, 82-99% male in model outputs) and cleaner (46% female in BLS, 80-92% male in outputs). Model generations do not improve steadily: bias worsens from SD 1.5 to SDXL before partially recovering in SD 3 Medium. A preliminary comparison with GPT-image-1 on five occupations suggests lower bias than open-source models, though the practical effect is small (Cramer's V = 0.080) and the comparison is exploratory. No model achieves gender parity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑