谁的安全观?面向文本到图像模型多元对齐的 Deep DIVE 数据集
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
- Google DeepMind(谷歌DeepMind)
- Google Research(谷歌研究)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出首个面向文本到图像模型多元对齐的多模态数据集 Deep DIVE,通过人口统计学交叉评估者收集多样化安全反馈,实证揭示伤害感知的群体差异,并探讨了构建更公平可引导 T2I 系统的策略。
AI中文摘要:
当前的文本到图像(T2I)模型往往未能考虑多样化的人类经验,导致系统出现错位。我们倡导多元对齐,即人工智能理解并能够被引导向多样化且常常相互冲突的人类价值观。我们的工作为实现 T2I 模型中的这一目标提供了三项核心贡献。首先,我们引入了一个用于多样化交叉视觉评估(DIVE)的新数据集——这是首个用于多元对齐的多模态数据集。它通过大量人口统计学交叉的人类评估者,在 1000 个提示上提供了广泛反馈,并具有高复现性,捕捉了细致的安全感知,从而实现对多样化安全视角的深度对齐。其次,我们通过实证确认了人口统计学特征是该领域中多样化观点的关键代理变量,揭示了与常规评估不同的、显著的、依赖情境的伤害感知差异。最后,我们讨论了构建对齐 T2I 模型的启示,包括高效的数据收集策略、LLM 判断能力以及模型向多样化视角的可引导性。本研究为更公平、更对齐的 T2I 系统提供了基础性工具。内容警告:本文包含可能有害的敏感内容。
英文摘要:
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.