arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26574cs.CRcs.AIcs.LG

恢复、解码、再防护:针对编码型VLM越狱的防护无关型防御放大技术

Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks

Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出防护无关型的恢复-解码放大器,评估其对11种攻击集成的编码型VLM越狱防御效果,发现其存在安全-效用上限,贡献了放大器、集成评估方法及防御适用场景图谱。

中文摘要 AI 辅助

安全分类器(“防护器”)是视觉语言模型(VLM)主流的黑盒防御手段,但它们仅判断输入的表面形式而非含义:将有害请求重新编码为集合论、形式逻辑、稀有语言、代码或文本图像,就能绕过原本会拦截该请求的防护器——这就是“解码差距”。自然的解决方案是构建一种防护无关型的恢复-解码放大器,它能转录图像内容并将编码文本重述为明文有效载荷,再提交给防护器,这样任何现成的分类器都能筛查真实请求。我们构建了该放大器并针对攻击者的最优场景进行评估:采用11种攻击的集成攻击,若任意攻击成功则判定行为被突破(遵循AutoAttack的最优套件设定,该指标在越狱防御中极少被报告),其突破率约为单攻击均值的3.5倍。这揭示了我们的核心发现:在5种防护器和2种目标VLM上,我们评估的非迭代恢复型防御存在经验性的安全-效用上限。该放大器仅部分缩小了差距:未受防护的集成攻击突破了89%-91%的行为,最优的防护器+放大器组合仍会留下63%-65%的漏洞;在10种防护器-目标模型配对中,仅4种配对的防护器性能提升具有统计显著性。该放大器在接口层面与防护器无关,但在效果层面并非如此。模块化的再防护层能大幅缩小剩余差距,但会使校准良好的防护器的良性请求弃权率升至81%-92%;唯一保持可用性的宽松防护器始终未达到可部署的安全标准(集成攻击成功率为48%)。针对所研究的流程以及表征转移攻击(即留下可识别有效载荷的编码和跨模态渲染,而非像素或嵌入空间攻击),我们评估的所有配置都无法同时实现低攻击成功率和低弃权率。我们贡献了该放大器、可凸显权衡关系的集成评估方法,以及基于恢复的VLM防御的适用场景与失效场景图谱。

英文摘要

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven encoding attacks - six published implementations, one standard encoding baseline, one adapted and three author-constructed renders - counting a behavior as broken if any attack succeeds. Restoring a view the guard never had is what buys coverage - block rates on image renders go from exactly zero to 67-90% - and what it costs in benign traffic is set by the guard, not by the mechanism: one guard pays 9 benign blocking points for the same 70-point gain another pays 69 for. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast for that step survives multiple-comparison correction. What does lower ensemble attack success is a reguard step that re-screens the recovered pre-decode surface, and it is the one every guard pays for: it raises benign over-refusal on all ten guard-target pairs, where restoring a single channel raises it on some and not others. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. That empty region is a property of the configurations we sample, not a bound on what recovery-based defenses can reach, and we breach its safety half ourselves.

补充信息

↑