arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00973cs.CL

关注差距:基于文本到图像系统中过滤器-生成器差异的零查询越狱攻击

Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems

Wanguang Li, Zhaoxin Wang, Handing Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对文本到图像系统的零查询越狱问题,提出基于过滤器-生成器差异的框架,通过筛选与集成进化搜索提升攻击成功率,在六个黑盒管线及商业服务上效果优于基线。

中文摘要 AI 辅助

文本到图像(T2I)系统通常在生成器前设置提示级安全过滤器以拦截不安全请求,但此类系统仍易受恶意越狱提示攻击。基于迁移的攻击可在不查询目标的情况下离线构建对抗提示,但往往过拟合于单个替代模型,且在庞大搜索空间中仅靠语义或感知相似性无法同时保证过滤器规避与保留不安全生成意图,浪费大量低潜力候选。我们发现过滤器与生成器以不同目标和表征处理同一提示,将此差距称为过滤器-生成器差异(FGD),该差异允许通过扰动降低提示对过滤器的感知风险,同时保留生成器所需的视觉概念。基于FGD,我们提出零查询越狱框架:在分词和语义阶段通过可观测差异规则将扰动筛选为高潜力候选集,再执行无需访问目标的替代集成进化搜索。对六个黑盒管线及一项商业在线服务的实验表明,我们的方法在六个管线中平均攻击成功率分别提升至29.2%(MHSC)和33.3%(Q16),较最强基线分别提升约8和12个百分点。

英文摘要

Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial prompts offline without querying the target, but they tend to overfit to a single surrogate. Moreover, they explore a large search space in which semantic or perceptual similarity alone cannot guarantee both filter evasion and preservation of the unsafe generation intent, wasting effort on low-potential candidates. We observe that the filter and the generator process the same prompt under different objectives and representations, and term this gap the Filter-Generator Discrepancy (FGD), which allows a perturbation to reduce a prompt's perceived risk to the filter while preserving the visual concept needed by the generator. Building on FGD, we propose a zero-query jailbreak framework that screens perturbations into a high-potential candidate set via observable discrepancy rules at the tokenization and semantic stages, and then performs a surrogate-ensemble evolutionary search that requires no access to the target. Experiments on six black-box pipelines and a commercial online service show that our method consistently outperforms representative baselines, raising the average attack success rate to 29.2\% (MHSC) and 33.3\% (Q16) across the six pipelines and improving over the strongest baseline by about 8 and 12 percentage points, respectively.

↑