发表机构
Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University; College of Artificial Intelligence, Tsinghua University; Security Department, Alibaba Group; School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(北京航空航天大学虚拟现实技术与系统国家重点实验室人工智能研究院; 清华大学人工智能学院; 阿里巴巴集团安全部; 中山大学深圳校区网络科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像生成易受隐式性提示影响的问题,提出UniNDM统一噪声驱动框架。利用早期预测噪声的可分离性开发轻量级检测器,引入噪声增强自适应负引导缓解问题,扩展到扩散变压器架构,实验显示比现有方法有显著改进。
AI 中文摘要
尽管文本到图像扩散模型具有强大的生成能力,但它们容易受到隐式性提示的影响,由于模型偏差或训练数据中的潜在相关性,微妙线索会意外生成不当内容。现有安全机制存在根本局限性。为此,我们提出UniNDM,一个统一的噪声驱动框架,通过扩散过程中的噪声动态来重新思考安全机制。我们发现早期预测噪声在正常和性明确内容之间具有内在可分离性,并理论证明其语义浓度随时间步长二次增加。利用此特性,我们开发了轻量级基于噪声的检测器,准确率高且几乎无计算开销。对于缓解,我们引入噪声增强自适应负引导,通过大语言模型动态生成特定上下文负提示,同时通过抑制对明确令牌的注意力集中来优化初始噪声。我们还将框架扩展到新兴的扩散变压器架构。综合实验表明,我们的方法比现有方法有显著改进。
英文摘要
Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or adversarial tokens unexpectedly generate the inappropriate content due to model biases or latent correlations in training data. Existing safety mechanisms face fundamental limitations: detection methods primarily identify explicit content and fail to capture implicit malicious intent, while mitigation approaches rely on static negative prompts inadequate for diverse implicit scenarios. To address these challenges, we propose UniNDM, a unified noise-driven framework that rethinks safety mechanisms through the lens of noise dynamics in diffusion processes. Our key insight is that early-stage predicted noise exhibits inherent separability between normal and sexually explicit content, which we theoretically demonstrates quadratically increasing semantic concentration with timestep. Leveraging this property, we develop a lightweight noise-based detector achieving superior accuracy with virtually no computational overhead. For mitigation, we introduce noise-enhanced adaptive negative guidance: dynamically generating context-specific negative prompts via large language models to handle diverse implicit content, while optimizing initial noise by suppressing attention concentration on explicit tokens to provide comprehensive protection. Besides the U-Net-based diffusion models, we further extend our framework to emerging Diffusion Transformer architectures through region-constrained semantic guidance tailored for their unified multimodal attention. Comprehensive experiments across U-Net models and DiT models on both natural and adversarial datasets demonstrate substantial improvements over state-of-the-art methods, including SLD, UCE, Safree, etc. Our code is publicly available at https://github.com/Aries-iai/UniNDM.
Comments18 pages, 10 figures, accepted by TPAMI