arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于安全文本到图像生成的内省注意力调制

Introspective Attention Modulation for Safe Text-to-Image Generation

Basim Azam, Hossein Rahmani, Naveed Akhtar

arXiv 2607.14945首次发表:更新:

发表机构

The University of Melbourne; Lancaster University(墨尔本大学; 兰卡斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于流的文本到图像模型易产生不安全内容的问题,提出通过推理时内省调节注意力动态实现安全的方法,该方法在保持或提升质量的同时显著提高安全分数,为安全图像生成提供新途径。

AI 中文摘要

基于流的最先进文本到图像(T2I)模型具有卓越的生成能力,但仍易产生不安全内容。先前的安全措施包括概念擦除、提示过滤和基于分类器的门控等。然而,像模型的参数高效调整等简单技术很容易绕过这些防护。我们引入一种独特的原则性方法,通过推理时的内省来调节模型的注意力动态以实现安全,具有内在鲁棒性。该方法在整个图像合成过程中分析并重新平衡注意力激活,在保持语义对齐的同时引导生成远离不安全概念。这种内省控制确保了部署模型的安全。在标准和对抗性安全基准测试中,我们的方法在保持甚至提高对齐和感知质量的同时,取得了显著的安全分数。我们的结果表明,与现有的概念擦除方法相比,注意力空间调节为基于扩散变压器的更安全图像生成提供了一条更有前景的途径。

英文摘要

State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Prior safety efforts range from concept erasure and prompt filtering to classifier-based gating. However, simple techniques like parameter efficient adaptations of the models easily bypass such guardrails. We introduce a unique principled approach that achieves safety by regulating the model's attention dynamics through inference-time introspection, exhibiting intrinsic robustness. Our method analyzes and rebalances attention activations throughout image synthesis, steering generations away from unsafe concepts while preserving semantic alignment. This introspective control ensures safety of deployed models. Across standard and adversarial safety benchmarks, our approach achieves remarkable safety scores while maintaining or even improving alignment and perceptual quality. Our results reveal that attention-space regulation offers a considerably more promising path to safer diffusion transformer based image generation than the existing concept erasing mechanism.Our code can be accessed at https://basim-azam.github.io/iam/

CommentsAccepted at ECCV 2026. 20 pages, 7 figures. Project page: https://basim-azam.github.io/iam/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑