发表机构
University of Zurich; University of Helsinki; ETH Zurich; Stanford University(苏黎世大学; 赫尔辛基大学; 苏黎世联邦理工学院; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出WORLDVIEW基准,通过审计提示词修订层,发现该层是文生图系统中文化偏见的因果来源,导致非西方语境被刻板化。
AI 中文摘要
商业文生图系统在生成图像之前会静默地修订用户提示词,这一步骤用户通常无法禁用甚至无法看到。然而,现有的文化偏见审计仅检查最终图像,并将生成过程视为单一流水线,因此无法确定偏见的来源。我们引入了WORLDVIEW,一个涵盖15种语言和31种语言-语境配对的多语言基准,包含8,960个提示词。利用该基准,我们通过三步分析审计了三个系统(DALL-E-3、Imagen-4、GPT-Image-1.5)中的修订层:它如何标记每种文化语境,是否将该语境压缩为狭窄的词汇,以及该词汇是否具有刻板印象。相对于无语境的英文基线,美国是最少被标记的语境,而非西方和非英语语境被标记得更为严重,被压缩为适用于主题多样提示词的狭窄词汇,并被简化为可识别的文化刻板印象。通过在没有修订层的模型上比较原始提示词与修订后提示词的图像,我们确定修订层本身是这种刻板印象的一个先前未被记录的因果来源。为了定位文化偏见并修复它,我们必须审计系统在实际部署中的表现,而不仅仅是模型本身。
英文摘要
Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell where the bias originates. We introduce WORLDVIEW, a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings. Using it, we audit the revision layer in three systems (DALL-E-3, Imagen-4, GPT-Image-1.5) through a three-step analysis of how heavily it marks each cultural context, whether it flattens that context into a narrow vocabulary, and whether that vocabulary is stereotypical. Relative to a no-context English baseline, the US is the least-marked context, while non-Western and non-Anglophone contexts are marked far more heavily, flattened into narrow vocabularies applied across topically diverse prompts, and reduced to recognizable cultural stereotypes. Comparing images from original versus revised prompts on models without a revision layer, we identify the layer itself as a previously undocumented, causal source of this stereotyping. To locate cultural bias, and fix it, we must audit the system as deployed, not the model alone.
CommentsAccepted to EMNLP 2026