arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于原子越狱策略解耦与组合的多模态自动红队评估

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination

Shiji Zhao, Yuxuan Zhou, Chen Xiong, Dongxian Wu, Yang Bai, Xun Chen

arXiv 2608.04034首次发表:更新:

AI 中文总结

针对现有多模态越狱攻击的两大挑战,本文提出HACA方法,构建多模态原子越狱策略空间,对5种主流MLLMs平均攻击成功率达95.48%。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)在图文理解与生成方面已取得显著进展,但仍易受到越狱攻击,触发有害输出,引发严重安全隐患。现有多模态越狱攻击虽已证明其可行性,却面临两大核心挑战:缺乏原子化多模态策略空间,以及缺少超越人工经验的简洁高效可执行框架。为解决上述挑战,本文首先将图文越狱策略空间分解为结构、语义、句法三个层级,构建涵盖文本与图像模态的越狱策略集,以系统实现对不同攻击类型的组合覆盖;随后提出名为分层原子组合攻击(Hierarchical Atomic Combination Attack, HACA)的多模态自动红队越狱方法,具体而言,基于六维策略空间,采用跨模态联合规划器选择并组合不同原子越狱策略,用于后续越狱指令生成;最后在执行层面,探索应用统一生成执行器,基于选定的多模态策略直接生成越狱指令。一系列实验表明,本文提出的自动红队方法对5种主流MLLMs的平均攻击成功率达95.48%。

英文摘要

Multimodal Large Language Models (MLLMs) have achieved impressive progress in image-text comprehension and generation, yet they remain susceptible to jailbreak attacks that can trigger harmful outputs and pose serious safety concerns. Existing multimodal jailbreak attacks have shown the feasibility of such attacks, but they still face two fundamental challenges: the lack of a atomic multi-modal strategy space, the absence of a concise and efficient executable framework beyond human-craft experience. To address these challenges, we first decompose the text-image jailbreak strategy space into three levels: structural, semantic, and syntactic, constructing a jailbreak strategy set encompassing both text and image modalities to systematically achieve combined coverage of different attack types. Then we propose a multimodal automated red team jailbreak method named Hierarchical Atomic Combination Attack (HACA). Specifically, based on a six-dimensional strategy space, a cross-modal joint planner is used to select and combine the different atomic jailbreak strategy for subsequent jailbreak command generation. Finally, at the implementation level, we explore to apply a unified generate executor to directly generate jailbreak instructions based on the selected multi-modal strategies. A series of experiments show that our automated red team method can achieve an attack success rate of average 95.48\% against five mainstream MLLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑