arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过结构化后缀建模加速扩散语言模型

Accelerating Diffusion Language Models via Structured Suffix Modeling

Zifeng Cheng, Keda Li, Zhiwei Jiang, Cong Wang, Fei Shen, Qing Gu

arXiv 2608.23167首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University; National University of Singapore(南京大学计算机软件新技术国家重点实验室; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对扩散语言模型(DLM)并行解码的高计算开销问题,提出结构化后缀建模方法,将后缀分区域处理并融入前一步解码信息,无需训练且可与其他加速技术结合,能进一步加速DLM推理并提升性能,长序列场景下最高可实现72.81倍加速。

AI 中文摘要

扩散语言模型(Diffusion Language Models, DLMs)在单一生成步骤中对多个标记进行去噪,展现出强大的并行解码能力。然而,这种并行性带来了巨大的计算开销,因为每一步都需要与所有后缀标记进行交互。现有方法通常通过仅保留局部后缀窗口来替代完整后缀以降低该成本,但这些方法忽略了后缀区域间的结构异质性,且在每个时间步都用相同表示重新初始化后缀标记。为此,我们提出一种用于高效DLM推理的结构化后缀建模方法。具体而言,我们将后缀划分为局部、中间和尾部三个区域,并根据各区域的结构角色为其保留不同数量的后缀标记。此外,我们将前一步的解码结果融入当前步骤的后缀标记表示中,使其能在各生成步骤中携带不断演化的去噪信息。值得注意的是,我们的方法无需训练,且与并行解码策略、KV缓存等多种现有加速技术正交。在三个DLM的多个基准测试上的实验结果表明,我们的方法可进一步加速DLM推理,且在多数情况下能提升性能;尤其在长序列推理中,结合其他加速技术时,我们的方法可实现最高72.81倍的加速。我们的代码可在此URL获取。

英文摘要

Diffusion Language Models (DLMs) exhibit strong parallel decoding capabilities by denoising multiple tokens in a single generation step. However, this parallelism comes with substantial computational overhead, as each step requires interactions with all suffix tokens. Existing methods typically reduce this cost by retaining only a local suffix window as a substitute for the full suffix. Despite their effectiveness, these methods overlook the structural heterogeneity across suffix regions and re-initialize suffix tokens with identical representations at each timestep. To this end, we propose a structured suffix modeling method for efficient DLM inference. Specifically, we divide the suffix into three regions, i.e., the local, middle, and tail regions, and retain different numbers of suffix tokens in each region according to their structural roles. Moreover, we incorporate the decoding results from the previous step into the suffix token representations at the current step, allowing them to carry evolving denoising information across generation steps. Notably, our method is training-free and orthogonal to several existing acceleration techniques, such as parallel decoding strategies and KV cache. Empirical results across multiple benchmarks on three DLMs demonstrate that our method can further accelerate DLM inference and improve performance in most cases. In particular, in long-sequence inference, our method achieves up to a \(72.81\times\) speedup when combined with other acceleration techniques. Our code is available at https://github.com/zifengcheng/SSM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑