发表机构
Applied Artificial Intelligence Initiative (A2I2), Deakin University; Pennsylvania State University(迪肯大学应用人工智能计划(A2I2); 宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过能量景观视角统一解释扩散语言模型越狱攻击,提出三个互补的无需训练检测信号,在多种模型上验证其互补性,且逃避检测的攻击均未产生有害内容。
AI 中文摘要
现有的针对基于扩散的大语言模型(dLLMs)的攻击和防御方法针对特定漏洞,但缺乏一个统一的框架来解释攻击为何成功。我们提出一个框架,将安全对齐解释为塑造去噪能量景观:一个对齐良好的模型通过能量屏障将有害查询引导至安全输出,该屏障分隔了这两个区域。当前的越狱攻击归结为两种规避此屏障的策略:在初始化时模糊查询的安全倾向,或在轨迹中途干预以迫使去噪路径跨越能量屏障。基于这一视角以及掩码扩散模型在去噪过程中最小化动能的结果,我们推导出三个互补的、无需训练的检测信号:一个步骤0比率,在生成开始前从logit分布读取初始安全倾向;以及两个轨迹速度信号,在logit空间的互补子空间中跟踪动能。攻击要么在初始化时暴露其意图,要么在至少一个被监控的子空间中消耗动能以跨越屏障,因此这三个信号在能量预算上按构造覆盖彼此的盲区。在三个密集dLLMs(LLaDA-8B、LLaDA-1.5、Dream-7B)和一个稀疏混合专家dLLM(LLaDA-MoE-7B)上的评估证实了这种互补性。在已知攻击的压力测试中,每个逃避检测的配置也未能产生有害内容,这表明检测阈值和屏障跨越阈值难以分离。
英文摘要
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three dense dLLMs (LLaDA-8B, LLaDA-1.5, Dream-7B) and a sparse mixture-of-experts dLLM (LLaDA-MoE-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.
Comments27 pages, 10 figures
Journal refNeurIPS 2026