发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将偏见评估框架BBG适配图像生成,对比6种T2I模型在照片、分镜、漫画生成中的偏见表现,发现叙事视觉形式中偏见更易显现,强调需用多样视觉形式评估T2I系统。
AI 中文摘要
文本到图像(T2I)生成模型正越来越多地应用于媒体内容创作、教育等场景,引发了人们对其输出可能再现社会偏见的担忧。已有研究表明T2I模型存在社会偏见,但现有评估大多聚焦于照片生成任务,因此尚不清楚此类偏见是否以及如何在分镜、漫画等更具叙事性的视觉形式中显现,这类形式会在多个面板中呈现角色与事件。本研究通过将基于文本的偏见评估框架BBG适配到图像生成任务,对比6种T2I模型在照片、分镜、漫画生成中的偏见表现。结果显示,专有模型在照片生成中平均产生25.9%的偏见输出,分镜生成中偏见输出占比增加9.6个百分点,漫画生成中增加18.2个百分点;照片主要通过微妙视觉线索编码偏见,而分镜和漫画则通过事件顺序、角色定位、叙事结局及文本元素更明确地展现偏见。这些发现表明,在照片生成中不易察觉的偏见可能会在叙事视觉形式中显现,凸显了用照片以外的多样视觉形式评估T2I系统的重要性。
英文摘要
Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.
CommentsAccepted to GenAI4World Workshop at COLM 2026