arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07430cs.LGcs.AI

扩散大语言模型作为目标与对抗者:机制性安全漏洞利用

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Elena Dumitrescu, Gert Lek, Lydia Y. Chen, Jérémie Decouchant

AI总结:

本研究揭示扩散大语言模型(DLLMs)的安全机制漏洞,提出SN-Guided Diffusion黑盒越狱框架,实现高迁移攻击成功率且生成成本远低于现有方法。

AI中文摘要:

扩散大语言模型(DLLMs)用迭代并行去噪替代自回归下一个词元预测,但其内部安全机制仍鲜为人知。本研究将DLLMs同时作为目标与对抗者,揭示基于扩散的对齐机制存在的漏洞。我们首先发现,DLLMs的安全对齐仍较为稀疏,且可跨架构迁移:源自自回归前身模型的DLLMs会继承源模型相同的机制性安全特征,支持通过直接安全神经元映射与剪枝实施迁移攻击。自剪枝使LLaDA的攻击成功率(ASR)从2.6%升至73.8%,Dream的ASR从1.9%升至86.6%;而基于Qwen2.5的迁移剪枝使Dream的ASR从1.9%升至73.2%,Fast-dLLM的ASR从7.0%升至86.3%。基于上述发现,我们提出SN-Guided Diffusion,这是一种完全离线的黑盒越狱框架,通过加权安全神经元损失引导扩散过程远离安全触发区域,实现近乎完美的提示可分性(良性与越狱提示区分的AUROC=1.0)。在多个开源与闭源目标模型上,该方法对Llama-3-8B-Instruct的迁移ASR最高达77.1%,对Qwen2.5-7B-Instruct达86.9%,对Gemini-2.5-Flash-Lite达74.3%,且每个提示仅需20次生成回合。与现有越狱框架相比,该方法在具备竞争力迁移性的同时,生成成本低几个数量级,代码库可通过指定URL获取。

英文摘要:

Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment. We first show that safety alignment in DLLMs remains sparse and transferable across architectures. DLLMs initialized from autoregressive predecessors inherit the same mechanistic safety footprint as their source models, enabling transfer attacks via direct safety neuron mapping and pruning. Self-pruning increases attack success rates (ASR) from 2.6% to 73.8% on LLaDA and from 1.9% to 86.6% on Dream, while transfer pruning from Qwen2.5 increases ASR from 1.9% to 73.2% on Dream and from 7.0% to 86.3% on Fast-dLLM. Building on these findings, we introduce SN-Guided Diffusion, a fully offline black-box jailbreak framework that steers the diffusion process away from safety-triggering regions using a weighted safety neuron loss, which achieves near-perfect prompt separability (AUROC = 1.0 for benign-vs-jailbreak discrimination). Across multiple open and proprietary targets, our method achieves a transfer ASR of up to 77.1% on Llama-3-8B-Instruct, 86.9% on Qwen2.5-7B-Instruct, and 74.3% against Gemini-2.5-Flash-Lite, while requiring only 20 generation episodes per prompt. Compared to prior jailbreaking frameworks, our method achieves competitive transferability with orders-of-magnitude lower generation cost. Our codebase is available at https://github.com/ellyoana/sn-guided-diffusion.

↑