早鸟解码:利用可学习块大小和并行采样加速扩散大语言模型
Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling
浏览论文内容
中文总结 AI 辅助
提出早鸟解码框架,通过可学习分组和位置感知采样在置信度阈值前并行解码低熵token,无需修改预训练权重,显著提升扩散大语言模型推理吞吐量。
中文摘要 AI 辅助
扩散大语言模型(dLLMs)通过迭代去掩码提供了一种有前景的并行解码范式,作为自回归生成的替代方案。然而,dLLMs通常需要多个步骤才能让token置信度达到解码阈值,即使采用块级KV缓存,也会导致推理效率低下。为了加速dLLM推理,我们首次提出了一种“早鸟(EB)”解码框架,其动机源于观察到具有相似低熵的token倾向于聚集,并且可以在达到置信度阈值之前被更早地联合解码。具体而言,我们的EB-Decode框架整合了两个关键使能组件:(1)一个可学习网络,能够自适应地将具有相似不确定性的token分组为可变长度块,而不是依赖固定块大小;(2)一个位置感知采样器,学习在预测的可变长度块内使用更少的解码步骤并行去掩码token。这两个组件都是在不修改预训练dLLM权重的情况下开发的,因此可以在服务期间直接作为插件部署,且训练和推理开销可忽略不计。在三个模型和四个基准上的大量实验一致验证了我们的观察和EB-Decode的有效性,与普通解码方法相比实现了3.53-18.76倍的吞吐量提升,与最强基线Fast-dLLM相比吞吐量提升高达1.58倍,同时精度相当。
英文摘要
Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative unmasking. However, dLLMs typically require many steps before token confidence reaches the decoding threshold, resulting in inefficient inference even with block-wise KV caching. To accelerate dLLM inference, we for the first time propose an "early-bird (EB)" decoding framework, motivated by the observation that tokens with similarly low entropy tend to cluster and can be jointly decoded earlier, before reaching the confidence threshold. In particular, our EB-Decode framework integrates two key enablers: (1) a learnable network that adaptively groups tokens with similar uncertainty into variable-length blocks, rather than relying on fixed block sizes; (2) a position-aware sampler that learns to unmask tokens in parallel using fewer decoding steps within predicted variable-length blocks. Both components are developed without modifying pretrained dLLM weights and can therefore be directly deployed as plug-ins during serving, with negligible training and inference overhead. Extensive experiments across three models and four benchmarks consistently validate our observation and the effectiveness of EB-Decode, achieving 3.53-18.76$\times$ higher throughput than the vanilla decoding method and up to 1.58$\times$ higher throughput over the strongest baseline, Fast-dLLM, with comparable accuracy.
发表机构
- Harvard University(哈佛大学)
- Georgia Institute of Technology(佐治亚理工学院)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。