发表机构
Technical University of Munich; Munich Center for Machine Learning (MCML); Munich Data Science Institute (MDSI); University of Trento; Orreco(慕尼黑工业大学; 慕尼黑机器学习中心; 慕尼黑数据科学研究所; 特伦托大学; 奥雷科公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ContextBias框架与ContextBench基准,评估文本到图像模型在角色关联视觉表征随上下文变化时的偏见持续性,发现语义不相关上下文不会抑制角色关联属性,凸显偏见基准需可控上下文变化的必要性。
AI 中文摘要
文本到图像模型会学习概念之间的关联,在本文中即人们的职业(我们称之为角色)与视觉属性之间的关联,这些关联是许多已观察到的刻板偏见形式的基础。该领域一个关键的开放问题是,当处于职业角色的人的视觉表征被置于不同的提示上下文时,这些关联是稳定的还是会发生变化。我们引入了ContextBias,这是一个可控评估框架,以及ContextBench,这是一个涵盖92个角色和1656个语义可控提示的基准,旨在分离上下文变化对角色关联视觉表征的影响。在对四个最先进的模型生成的66240张图像进行评估后,我们发现将角色置于语义不相关的上下文并不会抑制角色关联属性;相反,跨角色属性集中度增加(合并BI值提升0.047)。人口统计线索、特征服装和角色特定工具在无上下文、相关上下文和不相关上下文条件下仍然高度普遍,并且对语义提示重新表述具有鲁棒性。场景构图和相机取景表现出最高的上下文敏感性。这些发现揭示了一种在无上下文评估中基本不可见的刻板印象持续性形式,强调了在偏见基准测试中需要可控的上下文变化。代码和数据集:this https URL,this https URL
英文摘要
Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open question in this area is whether these associations are stable or change when visual representations of people in professional roles are placed in different prompted contexts. We introduce ContextBias, a controlled evaluation framework, and ContextBench, a benchmark spanning 92 roles and 1,656 semantically controlled prompts, designed to isolate the effect of contextual variation on role-linked visual representations. Evaluating four state-of-the-art models on 66,240 generated images, we find that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$). Demographic cues, characteristic garments, and role-specific tools remain highly prevalent across context-free, related, and unrelated conditions, and are robust to semantic prompt reformulation. Scene composition and camera framing show the greatest context-sensitivity. These findings reveal a form of stereotypical persistence that remains largely invisible to context-free evaluations, highlighting the need for controlled contextual variation in bias benchmarking. Code and dataset: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias
CommentsAccepted to EMNLP 2026