arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

充分释放多模态攻击者的潜力:视觉语言模型的元自适应越狱攻击

Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models

Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong

arXiv 2608.27531首次发表:更新:

AI 中文总结

提出元自适应多模态越狱攻击(MAMJ),通过优化攻击策略提示与攻击者权重提升多模态越狱效果,在MM-SafetyBench上对多款前沿VLMs的攻击成功率显著优于基线,且可迁移至未见过的目标模型。

AI 中文摘要

大型视觉语言模型(VLMs)的安全性正受到多模态越狱攻击的日益严峻考验,但现有攻击在元层面大多仍保持静态:基于模板的攻击固定图像-文本布局,而迭代攻击仅以固定攻击策略和冻结的攻击者参数调整图像-文本内容。本文提出元自适应多模态越狱攻击(MAMJ),该方法从两个维度优化攻击者本身:控制攻击迭代的攻击策略提示(ASP)θ,以及决定攻击有效性的攻击者权重φ。在多模态攻击轨迹组中,基于大型语言模型(LLM)的评判器首先优化θ,随后利用组聚合的攻击成功率(ASR)奖励更新φ。在MM-SafetyBench基准上,MAMJ对GPT-4o、Gemini-3-Pro-Preview和Seed 2.0的ASR分别达到81.0%、78.9%和82.3%,较最强的样本级基准高出最多24.1个百分点。学习得到的攻击者(θ⋆,φ⋆)无需重新训练即可迁移至未见过的目标模型,且在代表性防御措施下仍保持有效。这些结果揭示了前沿VLMs存在易受元自适应越狱攻击的系统性漏洞,并为针对元层面攻击者的防御措施提供了研究动机。代码可在指定URL获取。

英文摘要

The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image-text layout, while iterative attacks adapt only the image-text content with fixed attack strategies and frozen attacker parameters. We propose Meta-Adaptive Multimodal Jailbreaking (MAMJ), which instead optimizes the attacker itself along two axes: an attack strategy prompt (ASP) governing attack iteration and attacker model weights determining attack effectiveness. Across groups of multimodal attack trajectories, an LLM-based critique first refines the ASP, after which group-aggregated attack success rate (ASR) rewards update those weights. On MM-SafetyBench, MAMJ achieves 81.0%, 78.9%, and 82.3% ASR against GPT-4o, Gemini-3-Pro-Preview, and Seed 2.0, respectively, outperforming the strongest sample-level baseline by up to 24.1 percentage points. The learned attacker, comprising the optimized ASP and attacker weights, also transfers without retraining to unseen victims and remains effective under representative defenses. These results reveal a systemic vulnerability of frontier VLMs to meta-adaptive jailbreaks and motivate defenses against meta-level adversaries. Code is available at https://github.com/Alibaba-VELLDEPTH/MetaJailbreak-VLM.

CommentsAccepted by EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑