无辜的面板,仇恨的故事:评估与检测多轮视觉故事生成中的仇恨意图
Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
浏览论文内容
中文总结 AI 辅助
本研究针对多轮视觉故事生成的组级仇恨意图问题,构建了专用评估数据集,发现现有审核系统存在漏检,提出互补防御方法,强调安全需适配视觉叙事的发展。
中文摘要 AI 辅助
图画书和漫画长期以来被用于传播仇恨叙事,因为它们即使是儿童也能轻松理解,臭名昭著的纳粹宣传图画书《毒蘑菇》就是例证。最近,Gemini和GPT-Image等前沿文本到图像(T2I)系统支持跨轮生成具有一致角色和场景的对话式内容,这使得仇恨视觉故事——即共同传达仇恨叙事的有序图像组——的生成变得廉价且可规模化。尽管已有研究探讨了T2I系统生成的仇恨内容,但这些研究聚焦于单张图像,很大程度上未探索组级仇恨含义。本研究旨在填补这一空白。具体而言,我们推出了包含330个多轮配置的\texttt{HatefulStoryPrompts},这些配置来自两种语言、三种视觉风格的55个仇恨故事,并对五个前沿模型进行了4950次尝试的评估。每个模型完成了超过80%的故事,最强模型的完成率达到99.0%。我们进一步在\texttt{HatefulVisualStory}(一个包含969个仇恨图像集和990个良性对照的人工标注数据集)上评估了现有审核系统,发现它们经常遗漏组级仇恨含义:专用安全模型的召回率最高为34.9%,而一个强大的视觉语言模型达到67.5%。最后,我们提出了互补的主动防御和后生成防御:一个交互感知监控器在仅提示会话中实现了97.3%的召回率,在用户提供第一张图像时达到92.6%;联合分析完整图像组的后生成方法达到80.2%。本研究表明,随着图像生成从孤立输出发展为连贯视觉叙事,安全措施必须相应演进,从单图像审核转向对交互和图像关系的状态推理。
英文摘要
Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}. Recently, frontier text-to-image (T2I) systems such as Gemini and GPT-Image have enabled conversational generation with consistent characters and scenes across turns, making hateful visual stories, namely ordered image groups that collectively convey hateful narratives, cheap and scalable to produce. Although prior work has studied hateful content generation by T2I systems, it focuses on individual images, leaving group-level hateful meaning largely unexplored. We aim to address the gap. Concretely, we introduce \texttt{HatefulStoryPrompts}, comprising 330 multi-turn configurations from 55 hateful stories across two languages and three visual styles, and evaluate five frontier models over 4,950 attempts. Every model completes over 80\% of the stories, with the strongest reaching 99.0\%. We further evaluate existing moderation systems on \texttt{HatefulVisualStory}, a human-labeled dataset of 969 hateful image sets and 990 benign controls, and find that they frequently miss group-level hateful meaning: dedicated safety models achieve at most 34.9\% recall, while a strong vision-language model reaches 67.5\%. Finally, we propose complementary proactive and post-generation defenses. An interaction-aware monitor achieves 97.3\% recall for prompt-only sessions and 92.6\% when the user supplies the first image, while post-generation methods jointly analyzing completed image groups reach 80.2\%. Our work shows that, as image generation evolves from isolated outputs to coherent visual narratives, safety must evolve accordingly, from per-image moderation to stateful reasoning over interactions and image relationships.
发表机构
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
- Hewlett Packard Enterprise(惠普企业)
机构由 AI 辅助整理,请以论文原文为准。