发表机构
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对MM-DiTs的安全风险,提出一种无调优的概念擦除方法,通过在MM-DiT中间模块构建引导向量注入模型,实现高效鲁棒的概念擦除,性能优于现有方法。
AI 中文摘要
多模态扩散Transformer(MM-DiTs)在文本到图像生成任务中展现出卓越性能,超越了传统基于U-Net的扩散模型。然而,其强大的生成能力也带来了严重的安全隐患,可能生成敏感或不当内容。现有的概念擦除方法虽旨在缓解此类风险,但多数需要修改模型参数,这类方法通常依赖特定架构,难以应用于已部署的大型模型。部分无调优方法在应用于先进的大规模MM-DiTs时面临挑战,原因在于MM-DiTs内嵌的知识、广泛的语义空间以及依赖上下文的文本编码器。为应对这些挑战,我们提出通过直接操作模型内部表示来擦除概念。我们通过深入分析MM-DiT各模块的生成作用,得到关键见解:文本条件下的语义表示在MM-DiT的中间模块中最为显著。基于此,我们从中间模块中提取待擦除的有害概念和期望的安全概念的表示,通过两者的差值构建引导向量,并将该单一向量注入连续的早期和中间模块。我们的方法仅作用于稀疏的文本分支令牌,并利用整流流的稳定采样轨迹,实现了高效的概念擦除,且开销极小,无需任何训练。在多个MM-DiT模型上开展的大量实验表明,我们的方法在擦除各类概念方面达到了最优性能,能够对最终输出实现有效控制,且对对抗攻击具有鲁棒性。
英文摘要
Multimodal Diffusion Transformers (MM-DiTs) have demonstrated remarkable text-to-image generation performance, surpassing traditional U-Net-based diffusion models. Nevertheless, their powerful generative capabilities also raise significant safety concerns, as they may generate sensitive or inappropriate content. While existing concept erasure methods aim to mitigate such risks, most require modifying model parameters, which are often architecture-specific and impractical for deployed larger models. Several tuning-free approaches face challenges when applied to advanced large-scale MM-DiTs due to their deeply embedded knowledge, broad semantic space, and context-dependent text encoders. To address these challenges, we propose to erase concepts by directly manipulating the model's internal representations. Our key insight, derived from an in-depth analysis of MM-DiT's block-wise generative roles, is that text-conditioned semantic representations are most salient in the middle blocks of MM-DiTs. Based on this, we extract representations of an unwanted concept and a desirable safe one from the middle block, construct a steering vector from their difference, and inject this single vector into consecutive early and middle blocks. By operating exclusively on the sparse text-branch tokens and leveraging the straight sampling trajectory of rectified flow, our method achieves effective concept erasure with negligible overhead and without any training. Extensive experiments across MM-DiT models demonstrate that our method achieves state-of-the-art performance in erasing diverse concepts, enables effective control over the final output, and remains robust to adversarial attacks.
CommentsAccepted to ACM MM 2026