发表机构
Zhejiang University; Hong Kong Polytechnic University(浙江大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对T2I模型的安全漏洞,提出CollageAttack黑盒越狱方法,通过空间文本组合实现高达86.0%的攻击成功率,揭示跨模态安全差距。
AI 中文摘要
文本到图像(T2I)模型在语言理解、图像内文本渲染和视觉组合方面有了显著改进,但其安全机制并不总能跟上这些能力的发展。这创造了一个跨模态攻击面,其中有害语义在序列化提示中可能不明显,却通过图像级组合显现出来。我们提出了CollageAttack,一种自动化的单提示黑盒越狱方法,通过结合上下文相关场景、场景基础的文本载体和空间分布的文本片段,将语义组装转移到图像平面。在多个开源和商业T2I模型上的实验表明,CollageAttack实现了高达86.0%的攻击成功率,在同一模型上比最强基线高出18.5个百分点,同时持续产生更有害的输出并保留源意图。我们进一步发现,分布式的文本片段可以在生成后重建预期语义,视觉组合比纯文本产生更强的沟通影响。这些结果揭示了一个跨模态安全差距,其中有害含义来自单独不太明确的元素的组合。
英文摘要
Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicuous in a serialized prompt yet emerge through image-level composition. We propose CollageAttack, an automated single-prompt black-box jailbreak that shifts semantic assembly into the image plane by combining context-relevant scenes, scene-grounded textual carriers, and spatially distributed text fragments. Experiments across multiple open-weight and commercial T2I models show that CollageAttack achieves attack success rates of up to 86.0%, outperforming the strongest baseline on the same model by 18.5 percentage points, while consistently producing more harmful outputs and preserving the source intent. We further find that distributed textual fragments can reconstruct the intended semantics after generation, with visual composition producing stronger communicative impact than text alone. These results reveal a cross-modal safety gap in which harmful meaning emerges from the composition of individually less explicit elements.
CommentsContains potentially unsafe text-to-image generation examples. Code is released publicly