arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TYPO:针对商业闭源图像生成模型的指令密集型视觉越狱

TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

Meng Xie, Li Zeng, Hangtao Zhang, Xianlong Wang, Ziqi Zhou, Pengpeng Qiao, Zhetao Li

arXiv 2607.24897首次发表:更新:

AI 中文总结

研究针对商业闭源图像生成模型安全漏洞,提出TYPO框架,通过自动生成对抗性排版提示,利用文本和视觉双通道策略空间及自适应组合搜索,有效实现指令密集型视觉越狱,性能优于其他攻击且成本低。

AI 中文摘要

近期的商业图像生成模型能生成带可读文本的高质量图像,备受关注。但研究发现其存在安全漏洞,即拒绝直接生成有害文本,却允许在生成图像中以文本形式呈现相同内容。本文引入指令密集型视觉越狱概念,提出TYPO框架。它通过自动生成对抗性排版提示利用安全漏洞,将提示生成分解为文本和视觉通道,经自适应组合搜索优化策略组合。实验表明,TYPO在平均ASR上比九种代表性越狱攻击高出50.2%,平均查询成本仅0.04美元。

英文摘要

Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we first show that this same capability also introduces a previously unreported safety vulnerability: these systems may refuse to generate harmful text directly, yet permit the same content when rendered as text within generated images, i.e., safety alignment does not reliably transfer from textual outputs to text embedded in images. In this paper, unlike existing visual jailbreaks against image-generation models, which primarily induce models to generate harmful visual objects or scenes, we introduce the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images. Such outputs can amplify harm because the rendered instructions can be readily read and widely spread. To instantiate this threat, we propose TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts, which covertly steer image-generation models to express harmful intent as highly legible, typographically structured text. Specifically, TYPO decomposes prompt generation into two channels: a textual channel for reframing the target intent, and a visual channel for specifying its presentation form. We formulate these two channels as a dual-channel textual-visual strategy space and optimize candidate strategy combinations through an adaptive combinatorial search. Extensive experiments across four commercial models (i.e., GPT-Image-2, Nano Banana Pro, Qwen-Image-2, and Seedream 5.0 Lite) show that TYPO substantially outperforms nine representative jailbreak attacks by 50.2% in ASR on average, while incurring an average query cost of only $0.04.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑