arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24228cs.CV

IMPLICIT-Bench:在中性提示下测量文本到图像模型中的隐性偏见

IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts

Yue Dai, Ziyang Liu, Marc Cheong, Caren Han

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出IMPLICIT-Bench基准,通过知识图谱构建中性、刻板与反刻板提示三元组,测量T2I模型在中性提示下的隐性偏见,并揭示去偏见与语义保真度间的权衡。

中文摘要 AI 辅助

文本到图像(T2I)模型通常使用基于槽的模板(如“一张[职业]的照片”)来评估偏见。此类模板仅孤立地探测显性的人口统计属性(如性别、肤色),却忽略了自然提示中出现的更广泛的隐性偏见:当与刻板印象相关的属性未被指定时,模型仍会默认生成刻板印象的输出。我们引入了IMPLICIT-Bench,一个用于在此类提示下测量T2I模型中隐性偏见的基准。其关键设计是基于结构化知识图谱(KG)构建受控的提示三元组:中性、刻板印象和反刻板印象变体,这些变体仅沿单一偏见维度不同,同时保持场景语义不变。这使得能够精确归因偏见效应,而模板基准无法实现这一点。IMPLICIT-Bench包含11个偏见类别下的5,493个提示,并通过多模型一致性、基于CLIP的验证和人工评估进行了验证。利用该基准,我们表明最先进的T2I模型在中性提示下表现出系统性偏见,这是一种现有评估大多无法发现的失败模式。随后,我们使用IMPLICIT-Bench评估去偏见方法,揭示了偏见减少与语义保真度之间的根本权衡。

英文摘要

Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural prompts: when stereotype-relevant attributes are left unspecified, models still default to stereotypical outputs. We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts. The key design is a structured-knowledge-graph (KG) construction of controlled prompt triplets: neutral, stereotype, and anti-stereotype variants that differ only along a single bias dimension while preserving scene semantics. This enables precise attribution of bias effects that template benchmarks cannot achieve. IMPLICIT-Bench comprises 5,493 prompts across 11 bias categories, validated through multi-model agreement, CLIP-based verification, and human evaluation. Using this benchmark, we show that state-of-the-art T2I models exhibit systematic bias under neutral prompts, a failure mode largely invisible to existing evaluations. We then use IMPLICIT-Bench to evaluate debiasing methods, uncovering a fundamental trade-off between bias reduction and semantic fidelity.

发表机构

  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

↑