arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DC-Leap:通过草稿引导的连续跳跃解码实现无训练的dLLMs加速

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding

Yanhua Jiao, Tianyi Wu, Xiaoxi Sun, Yulin Li, HuiLing Zhen, Libo Qin, Baotian Hu, Zhuotao Tian, Min Zhang

arXiv 2607.20467首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Shenzhen Loop Area Institute; Huawei Noah’s Ark Lab(哈尔滨工业大学(深圳); 深圳河套学院; 华为诺亚方舟实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对dLLMs并行解码中因保守置信阈值致效率低的问题,提出DC-Leap框架。其采用动态连续验证策略消除JPDE,结合草稿引导解码机制。实验表明该框架能大幅加速dLLMs,长序列生成加速显著,且保证生成质量。

AI 中文摘要

虽然并行解码是扩散大语言模型(dLLMs)效率的核心,但当前策略常受过于保守的置信阈值阻碍。由联合概率依赖误差(JPDE)导致的这些阈值,造成冗余去噪迭代和次优推理速度。为克服此问题,我们提出DC-Leap,一个无训练框架,能在中等置信度范围内可靠加速dLLMs。DC-Leap引入动态连续验证策略,将严格有序的因果约束集成到并行解码过程中。通过逐步验证令牌依赖,该机制有效消除JPDE,实现可靠加速且性能相当。此外,DC-Leap纳入草稿引导解码机制,草稿通过跨多个令牌向前跳跃帮助扩展上下文,在推理时提供前瞻性上下文并保留双向注意力的结构优势。在标准基准上的大量实验表明,DC-Leap实现大幅加速,长序列生成在MBPP上高达53.19倍,与KV-Cache结合时高达105.02倍,且生成质量相当。代码可在给定网址获取。

英文摘要

While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime. DC-Leap introduces a Dynamic Contiguous Verification strategy that integrates strictly-ordered causal constraints into the parallel decoding process. By progressively validating token dependencies, this mechanism effectively neutralizes the JPDE, enabling reliable acceleration with comparable performance. Furthermore, DC-Leap incorporates the draft-guided decoding mechanism, where the draft helps extend the context by leaping forward across multiple tokens, providing look-ahead context and retaining the structural benefits of bidirectional attention during inference. Extensive experiments on standard benchmarks demonstrate that DC-Leap achieves substantial speedups, up to 53.19x on MBPP for long-sequence generation, and up to 105.02x when combined with KV-Cache with comparable generation quality. Code is available at https://github.com/ffh-wyls/DC-Leap .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑