arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态防御剖析实现文本到图像模型的认知越狱

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

Dongdong Yang, Deyue Zhang, Zhao Liu, Zonghao Ying, Wenzhuo Xu, Jiankai Jin, Xiangzheng Zhang, Quanchen Zou

arXiv 2607.17779首次发表:更新:

发表机构

Alibaba Tongyi Lab(阿里巴巴通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本到图像模型易受对抗性滥用问题,提出MIND认知越狱框架,将对抗提示生成重构为信念状态推理,集成多模态判断器等组件,经实验验证其在多个防御设置下显著优于现有方法,有效实现越狱生成。

AI 中文摘要

文本到图像(T2I)生成模型在合成高质量视觉内容方面取得了显著进展,但仍容易受到对抗性滥用,特别是在生成不适宜工作的(NSFW)图像方面。大多数现有越狱攻击主要依赖启发式提示工程或黑箱优化,将模型反馈视为二元信号(成功或失败)。这种粗粒度范式忽略了各种失败模式中嵌入的丰富信息,导致探索效率低下和严重的语义崩溃。本文提出了MIND,一个认知越狱框架,将对抗性提示生成重新构建为对潜在防御机制的信念状态推理问题。MIND通过将多模态反馈解释为高密度信号,积极地对目标系统的潜在防御机制进行建模。具体来说,该框架集成了三个核心组件:(1)用于细粒度反馈分解的多模态判断器,(2)用于迭代信念更新的防御剖析器,以及(3)用于检索历史有效攻击策略的元记忆模块。这些组件在一个推理驱动的进化优化过程中统一起来,实现自适应和语义一致的越狱生成。在I2P基准上的大量实验证明了MIND的有效性。在应用于Stable Diffusion v1.5模型的六种代表性预处理和后处理防御设置下,MIND实现了95.62%的攻击成功率(ASR),显著优于现有方法。此外,所提出框架的有效性在四个广泛使用的商业T2I系统中得到验证,在Wan-2.5上实现了91.58%的最高ASR。

英文摘要

Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existing jailbreak attacks primarily rely on heuristic prompt engineering or black-box optimization, treating model feedback as a binary signal (success or failure). This coarse-grained paradigm overlooks the rich information embedded in diverse failure modes, such as textual refusal, visual blocking, and semantic sanitization, resulting in inefficient exploration and severe semantic collapse. In this paper, we propose MIND, a cognitive jailbreak framework that reframes adversarial prompt generation as a belief-state inference problem over latent defense mechanisms. Instead of blindly searching for bypass prompts, MIND actively models the target system's latent defense mechanisms by interpreting multi-modal feedback as high-density signals. Specifically, the framework integrates three core components: (1) a Multi-modal Judge for fine-grained feedback decomposition, (2) a Defense Profiler for iterative belief updating, and (3) a Meta-Memory module for retrieving historically effective attack strategies. These components are unified within a reasoning-driven evolutionary optimization process, enabling adaptive and semantically consistent jailbreak generation. Extensive experiments on the I2P benchmark demonstrate the effectiveness of MIND. Under six representative pre-processing and post-processing defense settings applied to the Stable Diffusion v1.5 model, MIND achieves an Attack Success Rate (ASR) of 95.62%, significantly outperforming existing methods. Additionally, the effectiveness of the proposed framework is validated across four widely used commercial T2I systems, achieving the highest ASR of 91.58% on Wan-2.5.

Comments12 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑