DIVA:利用跨步条件传播实现离散扩散视觉语言模型的视觉越狱
DIVA: Exploiting Cross-Step Conditional Propagation for Visual Jailbreaks in Discrete Diffusion Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对多模态离散扩散视觉语言模型,提出跨步条件传播漏洞及DIVA白盒越狱框架,通过跨模态意图混淆和多时间步对抗优化,在三个模型上显著提升攻击成功率。
中文摘要 AI 辅助
大型视觉语言模型(VLMs)越来越多地部署在安全关键场景中,然而现有的视觉越狱研究几乎只关注自回归架构,留下了一个重要的新兴家族未被研究:多模态离散扩散视觉语言模型(dVLMs)。我们识别出一个扩散生成特有的漏洞:由于视觉嵌入条件作用于每一个反向去噪步骤,而非作为一次性前缀,对抗性视觉语义会在生成轨迹中反复传播和放大,我们将这一现象称为跨步条件传播。我们通过阶段敏感性分析、提示级切换率和成对去噪区间不一致性指标提供了经验证据,并通过自助重采样加以确认。我们提出了DIVA(离散扩散视觉语言模型攻击),这是一个白盒视觉越狱框架,采用跨模态意图混淆和扩散感知的多时间步对抗优化。在三个dVLMs上,DIVA在Beaver奖励模型指标下分别达到58.8%、67.7%和69.1%的HADES攻击成功率,优于为自回归模型设计的视觉越狱基线。代码:此https URL
英文摘要
Large vision-language models (VLMs) are increasingly deployed in safety-critical settings, yet existing visual jailbreak research has focused almost exclusively on autoregressive architectures, leaving an important emerging family unstudied: multimodal discrete diffusion vision-language models (dVLMs). We identify a vulnerability specific to diffusion generation: because the visual embedding conditions every reverse denoising step rather than acting as a one-time prefix, adversarial visual semantics are repeatedly propagated and amplified across the generation trajectory, a phenomenon we term cross-step conditional propagation. We provide empirical evidence through stage-sensitivity analysis, prompt-level switch rates, and pairwise denoising-bin disagreement metrics, confirmed by bootstrap resampling. We propose DIVA (Discrete-diffusion Vision-language model Attack), a white-box visual jailbreak framework using cross-modal intent obfuscation and diffusion-aware multi-timestep adversarial optimization. Across three dVLMs, DIVA reaches 58.8%, 67.7%, and 69.1% HADES ASR under the Beaver reward-model metric, outperforming visual jailbreak baselines designed for autoregressive models. Code: https://github.com/loststars2002/DIVA
发表机构
- Tsinghua University(清华大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。