发表机构
Seoul National University; Sungkyunkwan University(首尔大学; 成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散大语言模型可撤销解码中不可靠token损坏验证上下文的问题,提出无需训练的依赖感知可撤销解码框架,在12个基准上相比Saber实现2.71倍加速与4.35分CIDEr提升,优化了速度-质量权衡。
AI 中文摘要
扩散大语言模型(dLLMs)通过迭代去噪并行解码多个token,为自回归生成提供了有前景的替代方案。然而,增加解码并行度往往会降低生成质量,因为早期错误会污染后续上下文。可撤销解码通过重新评估已解码token并重新掩盖不可靠token来缓解该问题,但现有方法忽略了不可靠token也可能损坏验证上下文本身。我们识别出这种失效模式,提出了依赖感知可撤销解码(DARD),这是一种无需训练的框架,将token分为掩盖、候选和未掩盖状态。DARD使用排除不可靠token的选择性上下文来验证候选token,并自适应调节它们对后续解码的影响。在3个开源dLLM的12个文本和多模态基准上的实验表明,DARD相比近期的可撤销解码方法持续提升了速度-质量帕累托前沿,在Flickr30K上相比Saber实现了2.71倍加速和4.35分的CIDEr分数提升。
英文摘要
Diffusion large language models (dLLMs) offer a promising alternative to autoregressive generation by decoding multiple tokens in parallel through iterative denoising. However, increasing decoding parallelism often degrades generation quality, as early errors can contaminate later contexts. Revocable decoding mitigates this issue by re-evaluating decoded tokens and remasking unreliable ones, but existing methods overlook that unreliable tokens may also corrupt the verification context itself. We identify this failure mode and propose Dependency-Aware Revocable Decoding (DARD), a training-free framework that separates tokens into masked, candidate, and unmasked states. DARD verifies candidate tokens using a selective context that excludes less reliable tokens and adaptively regulates their influence on subsequent decoding. Experiments across 12 textual and multimodal benchmarks on 3 open-source dLLMs show that DARD consistently improves the speed-quality Pareto frontier over recent revocable decoding methods, achieving a 2.71$\times$ speedup and a 4.35-point CIDEr score gain over Saber on Flickr30K.