arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMJailBench:用于解耦多模态越狱漏洞的因子化基准

MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

Tianshi Wang, Jingsong Wang, Yafei Huang, Fengling Li, Xin Li, Lei Zhu

arXiv 2608.25490首次发表:更新:

发表机构

Tongji University; Mohamed bin Zayed University of Artificial Intelligence; Shanghai Artificial Intelligence Laboratory(同济大学; 穆罕默德·本·扎耶德人工智能大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员推出MMJailBench因子化基准,对16个MLLMs评估后揭示其越狱漏洞特征异质且依赖模型,明确提示框架等关键影响因素,开发出模块化多模态越狱评估套件用于高效审计。

AI 中文摘要

多模态大语言模型(MLLMs)正越来越多地被部署到实际应用中,但不同因素如何影响它们的越狱漏洞却鲜为人知。现有基准通常在单个越狱实例中结合了有害意图、提示框架、视觉语义和指令载体,掩盖了观察到的漏洞的具体来源。为解决这一局限,我们推出了MMJailBench,这是一个因子化基准,可在受控配置下系统地改变和组合这些因素,实现细粒度比较和因子级归因。对16个开放权重和专有MLLMs的大规模评估显示,漏洞特征具有高度异质性且依赖于模型。越狱漏洞在不同危害领域差异显著,暴露出当前多模态安全对齐的覆盖不均衡。提示框架是差异的主要来源,与任务相关的视觉语义会系统性地增加越狱易感性,其中权威类线索会暴露出特别明显的漏洞,而视觉呈现的指令相比直接文本指令并不会持续增加越狱易感性。为进一步研究多模态语境带来的风险,我们对一个代表性开放权重模型进行了诊断分析,在内部表征和跨模态交互中识别出与漏洞相关的模式。最后,我们开发了一个模块化多模态越狱评估套件,具有完整和轻量配置、多种评判选项以及多维指标,支持可复现、可扩展且成本效益高的多模态越狱审计。

英文摘要

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within individual jailbreak instances, obscuring the specific sources of observed vulnerabilities. To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and combines these factors under controlled configurations, enabling fine-grained comparison and factor-level attribution. Large-scale evaluations across 16 open-weight and proprietary MLLMs reveal highly heterogeneous and model-dependent vulnerability profiles. Jailbreak vulnerability varies markedly across harm domains, exposing uneven coverage in current multimodal safety alignment. Prompt framing emerges as the dominant source of variation, task-relevant visual semantics systematically increase jailbreak susceptibility with authority-like cues exposing particularly pronounced vulnerabilities, and visually rendered instructions do not consistently increase jailbreak susceptibility relative to direct textual instructions. To further investigate the risks introduced by multimodal context, we conduct diagnostic analyses on a representative open-weight model and identify vulnerability-associated patterns in internal representations and cross-modal interactions. Finally, we develop a modular multimodal jailbreak evaluation suite with full and lightweight configurations, multiple judge options, and multidimensional metrics, enabling reproducible, scalable, and cost-efficient multimodal jailbreak auditing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑