arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

诱饵图像增强针对编码越狱的字幕介导防御

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita

arXiv 2608.01043首次发表:更新:

AI 中文总结

该研究发现附加诱饵图像可降低VLMs的编码越狱攻击成功率,通过编码输入检测器门控附加诱饵,可在保留安全增益的同时控制良性弃权率,该效应在多模型、多场景下可复现。

AI 中文摘要

我们报告了视觉语言模型(VLMs)上图像输入与现有黑盒防御之间的一种反直觉交互:将编码越狱提示与无关诱饵图像配对可大幅降低攻击成功率(ASR)。关键变化在于防御流程而非图像本身。在五个前沿VLMs、两类编码攻击家族和三类黑盒防御中,原本对纯文本编码输入ASR基本无影响的字幕介导防御(ECSO),在附加无内容诱饵后ASR降幅最高达73个百分点;所有非饱和对比经精确麦克尼马尔检验均显著。由于黑盒威胁模型禁止检查厂商内部,我们基于间接证据提出两个假设以解释该现象:字幕介导防御会根据图像存在与否分支,图像侧的内在安全性会因图像中存在内容而激活。三项对照实验约束了该解释:空白画布和自然照片诱饵在所有模型上均重现了该效应,说明是图像存在而非内容导致;该效应在三个无 moderation 层的开放权重VLMs上可复现,因此并非厂商过滤的人为结果;非符号化、基于意义的编码器也能重现该效应,故并非特定于符号混淆。无条件附加诱饵不可部署——它会将良性弃权(不执行)率提升至20%-79%,增幅为+10至+67个百分点——但通过轻量编码输入检测器来门控附加操作,可将良性弃权率恢复至文本基线,同时在检测器触发处保留安全增益,此时检测器召回率成为约束条件。在针对字幕介导重检查的自适应攻击下,该效应会减弱但仍存在。我们将此视为关于流程交互的观察,而非一种稳健防御。

英文摘要

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five frontier VLMs, two encoded-attack families, and three black-box defenses, a caption-mediated defense (ECSO) that leaves ASR essentially unchanged on text-only encoded input drops it by up to $73$pp once a content-free decoy is attached; every non-saturated contrast is significant under exact McNemar tests. We advance two hypotheses for this pattern, supported by indirect evidence rather than pipeline introspection, since a black-box threat model precludes inspecting vendor internals: caption-mediated defenses branch on image presence, and intrinsic image-side safety engages on image-resident content. Three controls constrain the explanation. Blank-canvas and natural-photograph decoys reproduce the effect on every model, implicating image presence rather than content; the effect replicates on three open-weight VLMs served with no moderation layer, so it is not a vendor-filtering artifact; and a non-symbolic, meaning-based encoder reproduces it, so it is not specific to symbolic obfuscation. Attaching a decoy unconditionally is not deployable --- it raises benign refusal to $20$--$79\%$, an inflation of $+10$ to $+67$pp --- but gating attachment on a lightweight encoded-input detector returns benign refusal to the text baseline while preserving the safety gain wherever the detector fires, making detector recall the binding constraint. Under adaptive attacks that target the caption-mediated re-check, the effect degrades but holds. We frame this as an observation about pipeline interaction, not as a robust defense.

CommentsThere is bug in algorithm implementation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑