发表机构
Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像模型生成NSFW内容的问题,提出一种生成过程中检测NSFW信号并结合强化学习引导的安全生成框架,在标准和对抗性评估中优于现有方法。
AI 中文摘要
最近的文本到图像(T2I)模型在视觉图像生成性能上表现出色,但它们仍然可能生成包含暴力或露骨内容的NSFW(不适合工作场所)内容。现有的安全检查机制主要局限于生成前过滤(例如基于提示词级别的文本分类器)或图像完全合成后的事后审核。然而,对抗性攻击方法在更广泛的空间中运作。这种不平衡凸显了在生成过程中进行干预的安全机制的必要性。我们提出了一种生成过程中的安全框架,该框架监控去噪轨迹并从中间表示中检测新兴的NSFW信号。我们的方法不仅仅是检测NSFW生成,而是应用强化学习从NSFW提示中生成安全图像。通过将生成过程中的检测与可控引导相结合,我们的方法即使在生成已经开始后出现NSFW信号时也能缓解不安全轨迹。实验结果表明,我们的方法在标准和对抗性评估集上均持续优于现有的安全图像生成方法,同时保持感知质量和提示保真度。代码将在论文被接收后发布。
英文摘要
Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.