arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17067cs.AI

DiSCO:通过分布引导的对比提示优化防御文本到图像生成

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem

首次发表
浏览论文内容

中文总结 AI 辅助

DiSCO是一种零样本黑盒防御模块,通过分布引导的对比提示优化,在I2P基准上分别将未防御、已防御模型的攻击成功率降低37.7%、25.13%,提升文本到图像生成的安全性且保持语义保真度。

中文摘要 AI 辅助

随着文本到图像生成模型的发展,它们引发了关键的安全问题,尤其是生成暴力、裸体等非工作安全(NSFW)内容,而红队对抗攻击进一步加剧了这一问题。现有的防御措施大多在白盒假设下运行,依赖文本编码器优化、权重编辑或推理时干预,且根本无法扩展到专有模型。基于大语言模型(LLM)提示重写的黑盒替代方案适用性更广,但在我们识别的“良性对抗”问题上失效:这类提示在语言层面是安全的,但由于模型学习到的数据分布仍会触发有害生成。我们提出DiSCO,这是一种零样本、严格的黑盒防御,完全在提示层面作为即插即用模块运行,无需模型重新训练、微调或访问模型内部结构。DiSCO通过束搜索执行分布引导的后缀扩展,利用目标模型自身生成的安全与有害图像池进行对比评分优化,迭代自适应反馈直至生成安全内容。我们在I2P基准上的多次红队攻击下,证明DiSCO能持续提升未防御和已防御模型的安全性,分别实现37.7%和25.13%的攻击成功率(ASR)降低,同时保持语义保真度并提升图像连贯性。作为黑盒、架构无关的模块,DiSCO可轻松应用于任何文本到图像系统,无需对模型本身做任何更改。

英文摘要

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \textit{benign adversarial} problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

发表机构

  • Qualcomm AI Research(高通人工智能研究院)
  • King Abdullah University of Science Technology (KAUST)(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑