发表机构
College of Computer Science and Technology, National University of Defense Technology; Shenzhen MSU-BIT University; City University of Hong Kong(国防科技大学计算机科学与技术学院; 深圳北理莫斯科大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CRACK多智能体辩论框架,通过统一几何框架分析文生图模型安全过滤器的跨层冲突,实现高攻击成功率且查询量少、语义保真度高的越狱攻击。
AI 中文摘要
尽管文生图(T2I)模型日益受到由文本过滤器、图像分类器和跨模态检测器组成的异构多层安全栈的保护,它们仍易遭受引发非工作安全(NSFW)内容的越狱攻击。现有越狱研究要么针对单个过滤器优化,要么用聚合反馈查询完整流水线,难以识别活跃约束并适应安全间的冲突。本文提出了Detection Surface,一个统一的几何框架,用于表征异构T2I安全过滤器诱导的决策边界及其对越狱搜索空间的联合影响。该公式表明,成功规避由跨层冲突形成的稀疏非凸区域决定,绕过一个过滤器的突变可能会增加对另一个过滤器的暴露。受此分析启发,我们提出了CRACK,一个用于自适应越狱搜索的多智能体辩论框架,将越狱搜索分解为探索、诊断和仲裁。CRACK协调攻击智能体、防御智能体和评判智能体,迭代生成提示突变、获取特定层的诊断反馈,并通过奖励引导的优化突变策略。通过多轮辩论,CRACK在保留原始有害意图的同时,调整其搜索方向以适应不断变化的跨层约束。在多个T2I模型、数据集和安全配置上的大量实验表明,CRACK在复合防御下达到高达99.63%的攻击成功率(ASR),同时比现有方法需要更少的查询并保持语义保真度。
英文摘要
Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and cross-modal detectors. Existing jailbreak studies either optimize against individual filters or query the complete pipeline with aggregate feedback, making it difficult to identify the active constraint and adapt to conflicts across safety layers. In this paper, we introduce the Detection Surface, a unified geometric framework that characterizes the decision boundaries induced by heterogeneous T2I safety filters and their joint effect on the jailbreak search space. This formulation reveals that successful evasion is governed by a sparse and non-convex region shaped by cross-layer conflicts, where mutations that bypass one filter may increase exposure to another. Motivated by this analysis, we propose CRACK, a multi-agent debate framework for adaptive jailbreak search that decomposes jailbreak search into exploration, diagnosis, and arbitration. CRACK coordinates an Attack Agent, a Defense Agent, and a Judge Agent to iteratively generate prompt mutations, obtain layer-specific diagnostic feedback, and optimize mutation strategies through reward-guided refinement. Through repeated rounds of debate, CRACK adapts its search direction to the evolving cross-layer constraints while preserving the original harmful intent. Extensive experiments across multiple T2I models, datasets, and safety configurations show that CRACK achieves Attack Success Rates (ASR) of up to 99.63% under composite defenses, while requiring fewer queries than existing methods and maintaining semantic fidelity.
Comments16 pages, 11 figures