AI 中文总结
提出元自适应多模态越狱攻击(MAMJ),通过优化攻击策略提示与攻击者权重提升多模态越狱效果,在MM-SafetyBench上对多款前沿VLMs的攻击成功率显著优于基线,且可迁移至未见过的目标模型。
AI 中文摘要
大型视觉语言模型(VLMs)的安全性正受到多模态越狱攻击的日益严峻考验,但现有攻击在元层面大多仍保持静态:基于模板的攻击固定图像-文本布局,而迭代攻击仅以固定攻击策略和冻结的攻击者参数调整图像-文本内容。本文提出元自适应多模态越狱攻击(MAMJ),该方法从两个维度优化攻击者本身:控制攻击迭代的攻击策略提示(ASP)θ,以及决定攻击有效性的攻击者权重φ。在多模态攻击轨迹组中,基于大型语言模型(LLM)的评判器首先优化θ,随后利用组聚合的攻击成功率(ASR)奖励更新φ。在MM-SafetyBench基准上,MAMJ对GPT-4o、Gemini-3-Pro-Preview和Seed 2.0的ASR分别达到81.0%、78.9%和82.3%,较最强的样本级基准高出最多24.1个百分点。学习得到的攻击者(θ⋆,φ⋆)无需重新训练即可迁移至未见过的目标模型,且在代表性防御措施下仍保持有效。这些结果揭示了前沿VLMs存在易受元自适应越狱攻击的系统性漏洞,并为针对元层面攻击者的防御措施提供了研究动机。代码可在指定URL获取。
英文摘要
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image-text layout, while iterative attacks adapt only the image-text content with fixed attack strategies and frozen attacker parameters. We propose Meta-Adaptive Multimodal Jailbreaking (MAMJ), which instead optimizes the attacker itself along two axes: an attack strategy prompt (ASP) governing attack iteration and attacker model weights determining attack effectiveness. Across groups of multimodal attack trajectories, an LLM-based critique first refines the ASP, after which group-aggregated attack success rate (ASR) rewards update those weights. On MM-SafetyBench, MAMJ achieves 81.0%, 78.9%, and 82.3% ASR against GPT-4o, Gemini-3-Pro-Preview, and Seed 2.0, respectively, outperforming the strongest sample-level baseline by up to 24.1 percentage points. The learned attacker, comprising the optimized ASP and attacker weights, also transfers without retraining to unseen victims and remains effective under representative defenses. These results reveal a systemic vulnerability of frontier VLMs to meta-adaptive jailbreaks and motivate defenses against meta-level adversaries. Code is available at https://github.com/Alibaba-VELLDEPTH/MetaJailbreak-VLM.
CommentsAccepted by EMNLP 2026 Main Conference