超越双向承诺:重新评估扩散语言模型的鲁棒性
Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
该研究评估了扩散语言模型的鲁棒性,发现其虽能抵御部分对抗攻击但易受自然噪声影响,存在过度自信问题,脆弱性源于解码器路由故障,需将鲁棒性整合入迭代解码循环。
中文摘要 AI 辅助
扩散语言模型(Diffusion Language Models, DLMs)通过支持双向上下文与迭代优化,为自回归(autoregressive, AR)生成提供了极具吸引力的替代方案,但其在自然输入噪声与对抗攻击下的可靠性仍未得到充分探索。为解决这一问题,我们针对DLM的鲁棒性与校准度开展系统评估,并与自回归基准模型进行对比,采用两组参数匹配的模型对(LLaDA-8B vs. LLaMA-3-8B、Dream-7B vs. Qwen2.5-7B),覆盖32种自然扰动条件、对抗梯度探测及机制性隐状态分析。这种配对设计可有效分离架构固有属性与权重依赖行为。我们发现DLM的鲁棒性特征较为复杂:高度随机的DLM损失景观可自然抵御基于梯度的对抗后缀,但无法保证抵御自然噪声,证明日常鲁棒性依赖于权重而非架构固有特性;此外,DLM存在系统性过度自信问题,构成实际部署隐患。最关键的是,机制性探测显示所有模型均能完美编码输入损坏,将行为脆弱性完全归因于解码器路由故障。基于该诊断,我们表明表层提示修补无法优于噪声基线,最终得出结论:DLM的鲁棒性无法通过修补实现,必须根本性地整合到迭代解码循环中。
英文摘要
Diffusion Language Models (DLMs) offer a compelling alternative to autoregressive (AR) generation by enabling bidirectional context and iterative refinement. However, their reliability under natural input noise and adversarial attacks remains under-explored. To address this, we systematically evaluate DLM robustness and calibration against AR baselines, using two parameter-matched pairs (LLaDA-8B vs. LLaMA-3-8B and Dream-7B vs. Qwen2.5-7B) across 32 natural perturbation conditions, adversarial gradient probes, and mechanistic hidden-state analyses. This paired design effectively isolates architecture-intrinsic properties from weight-dependent behaviors. We find a nuanced robustness profile: while highly stochastic DLM loss landscapes naturally resist gradient-based adversarial suffixes, they provide no guaranteed defense against natural noise, proving that everyday robustness is weight-dependent rather than inherently architectural. Furthermore, DLMs exhibit systematic overconfidence, presenting a practical deployment hazard. Most crucially, mechanistic probing reveals that all models perfectly encode input corruption, isolating behavioral fragility entirely to a decoder routing failure. Consistent with this diagnosis, we show that surface-level prompt patching fails to improve over noisy baselines. Ultimately, DLM robustness cannot be patched on; it must be fundamentally integrated into the iterative decoding loop.
发表机构
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。