恢复、解码、再防护:针对编码型VLM越狱的防护无关型防御放大技术
Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
浏览论文内容
中文总结 AI 辅助
本研究提出防护无关型的恢复-解码放大器,评估其对11种攻击集成的编码型VLM越狱防御效果,发现其存在安全-效用上限,贡献了放大器、集成评估方法及防御适用场景图谱。
中文摘要 AI 辅助
安全分类器(“防护器”)是视觉语言模型(VLM)主流的黑盒防御手段,但它们仅判断输入的表面形式而非含义:将有害请求重新编码为集合论、形式逻辑、稀有语言、代码或文本图像,就能绕过原本会拦截该请求的防护器——这就是“解码差距”。自然的解决方案是构建一种防护无关型的恢复-解码放大器,它能转录图像内容并将编码文本重述为明文有效载荷,再提交给防护器,这样任何现成的分类器都能筛查真实请求。我们构建了该放大器并针对攻击者的最优场景进行评估:采用11种攻击的集成攻击,若任意攻击成功则判定行为被突破(遵循AutoAttack的最优套件设定,该指标在越狱防御中极少被报告),其突破率约为单攻击均值的3.5倍。这揭示了我们的核心发现:在5种防护器和2种目标VLM上,我们评估的非迭代恢复型防御存在经验性的安全-效用上限。该放大器仅部分缩小了差距:未受防护的集成攻击突破了89%-91%的行为,最优的防护器+放大器组合仍会留下63%-65%的漏洞;在10种防护器-目标模型配对中,仅4种配对的防护器性能提升具有统计显著性。该放大器在接口层面与防护器无关,但在效果层面并非如此。模块化的再防护层能大幅缩小剩余差距,但会使校准良好的防护器的良性请求弃权率升至81%-92%;唯一保持可用性的宽松防护器始终未达到可部署的安全标准(集成攻击成功率为48%)。针对所研究的流程以及表征转移攻击(即留下可识别有效载荷的编码和跨模态渲染,而非像素或嵌入空间攻击),我们评估的所有配置都无法同时实现低攻击成功率和低弃权率。我们贡献了该放大器、可凸显权衡关系的集成评估方法,以及基于恢复的VLM防御的适用场景与失效场景图谱。
英文摘要
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven encoding attacks - six published implementations, one standard encoding baseline, one adapted and three author-constructed renders - counting a behavior as broken if any attack succeeds. Restoring a view the guard never had is what buys coverage - block rates on image renders go from exactly zero to 67-90% - and what it costs in benign traffic is set by the guard, not by the mechanism: one guard pays 9 benign blocking points for the same 70-point gain another pays 69 for. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast for that step survives multiple-comparison correction. What does lower ensemble attack success is a reguard step that re-screens the recovered pre-decode surface, and it is the one every guard pays for: it raises benign over-refusal on all ten guard-target pairs, where restoring a single channel raises it on some and not others. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. That empty region is a property of the configurations we sample, not a bound on what recovery-based defenses can reach, and we breach its safety half ourselves.