arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本到图像模型中的有害内容生成:能力与审核限制

Harmful Content Generation in Text-to-Image Models: Capabilities and Moderation Limitations

Paschalis Giakoumoglou, Manos Schinas, Symeon Papadopoulos

arXiv 2610.06503首次发表:更新:

发表机构

Centre for Research and Technology Hellas (CERTH)(希腊研究与技术中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了五个开源文本到图像模型的有害内容生成能力,发现高生成率(如血腥89.2%),并揭示了审核与检测系统的显著漏洞,强调需加强防御。

AI 中文摘要

文本到图像生成模型能够产生高度逼真的图像,但也引发了对有害滥用的担忧。尽管存在安全机制,但针对现实攻击的有效性的系统性评估仍然有限。我们使用一个自动化流水线,将合法的新闻标题转化为针对色情内容、暴力/血腥、有害刻板印象、自残和仇恨言论的不安全提示,对五个开源文本到图像模型进行了有害内容生成的系统性评估。我们评估了具有内置安全机制的标准模型和绕过内容限制的社区微调变体。对1,500张生成图像的人工评估显示,有害内容生成率很高:血腥相关提示为89.2%,色情内容为47.6%,有害刻板印象为43.6%,仇恨言论为46.0%,自残为34.5%,主要通过图形暴力实现。模型在生成暴力和刻板印象内容方面表现出显著能力,而社区微调变体尤其容易受到色情提示的影响。在有害提示下,生成质量基本得以保留,产生的图像保真度足以构成虚假信息和滥用的风险;FLUX.1-dev在30.9%的情况下生成清晰逼真的有害图像。我们进一步评估了自动化审核系统,发现存在显著的检测漏洞,使得不安全图像能够逃避过滤。最后,我们评估了合成图像检测器,表明仅在良性数据集上训练的模型在显式内容上表现较差,而更多样化的训练数据提高了检测性能,凸显了当前方法中的语义分布差距。这些发现暴露了当前生成保障措施、审核系统和合成图像检测的局限性,强调了需要更强大的防御措施来应对大规模滥用。

英文摘要

Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic evaluation of harmful content generation across five open text-to-image models using an automated pipeline that transforms legitimate news captions into unsafe prompts targeting sexually explicit content, violence/gore, harmful stereotypes, self-harm, and hate speech. We evaluate both standard models with built-in safety mechanisms and community fine-tuned variants that bypass content restrictions. A human evaluation of 1,500 generated images shows high harmful-content generation rates: 89.2% for gore-related prompts, 47.6% for sexually explicit content, 43.6% for harmful stereotypes, 46.0% for hate speech, and 34.5% for self-harm, predominantly through graphic violence. Models show substantial capability for generating violent and stereotypical content, while community fine-tuned variants are particularly vulnerable to sexually explicit prompts. Generation quality is largely preserved under harmful prompting, producing imagery of sufficient fidelity to pose risks for disinformation and abuse; FLUX.1-dev produces clearly realistic harmful images in 30.9% of cases. We further evaluate automated moderation systems and find substantial detection gaps that allow unsafe images to evade filtering. Finally, we assess synthetic image detectors and show that models trained only on benign datasets perform worse on explicit content, while more diverse training data improves detection, highlighting semantic distribution gaps in current approaches. These findings expose limitations in current generation safeguards, moderation systems, and synthetic image detection, highlighting the need for stronger defenses against misuse at scale.

CommentsAccepted for publication in ACM Transactions on Intelligent Systems and Technology (TIST)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑