超越 token 位置:扩散语言模型中跨去噪步骤的安全对齐
Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对扩散大语言模型(dLLMs)的安全对齐问题,提出无需训练的解码方法 RAEC,通过利用早期去噪步骤的拒绝信号,在 LLaDA 和 Dream 上降低了攻击成功率且保留了实用性。
中文摘要 AI 辅助
扩散大语言模型(dLLMs)通过迭代去噪而非从左到右解码生成文本。这种生成范式引入了两个可影响安全对齐的维度:token 在去噪过程中何时生成,以及它们在响应中出现在何处。本文中,我们通过追踪去噪过程中的中间 token 分布和承诺决策,测量 dLLM 在有害提示下的安全行为。分析显示,拒绝信号集中在早期去噪步骤和响应的前导位置,早期承诺的 token 可强烈影响最终安全结果。我们的测量进一步表明,去噪步骤和拒绝 token 承诺的持续性对理解 dLLM 安全至关重要。基于这些发现,我们提出 Refusal-Aware Early Commitment(RAEC),一种无需训练的简单解码方法,用于从早期步骤中承诺持续的拒绝信号。在 LLaDA 和 Dream 上的实验表明,RAEC 可降低攻击成功率,同时在很大程度上保留实用性。代码可在该 https URL 获取。
英文摘要
Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the response. In this paper, we measure dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising. Our analysis shows that refusal signals are concentrated in early denoising steps and leading response positions, and the tokens committed early can strongly shape the final safety outcome. Our measurements further show that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Based on these findings, we propose Refusal-Aware Early Commitment (RAEC), a simple training-free decoding method that commits persistent refusal signals from early steps. Experiments on LLaDA and Dream show that RAEC reduces attack success rates while largely preserving utility. The code is available at https://github.com/Glresearch1/RAEC.
发表机构
- Case Western Reserve University(凯斯西储大学)
机构由 AI 辅助整理,请以论文原文为准。