arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26796cs.CL

Flash-dLLM:面向快速、内存高效的扩散大语言模型的IO感知KV缓存与并行解码

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

首次发表
浏览论文内容

中文总结 AI 辅助

Flash-dLLM提出IO感知融合KV缓存与自草拟验证并行解码,无需训练即加速扩散LLM推理,在GSM8K和HumanEval上分别比Elastic-Cache快5.1倍和11.0倍。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)近期作为自回归大语言模型的一种有前景的替代方案出现,通过支持非自回归文本生成。然而,其实际部署仍受限于推理效率低下,这主要是由于缺乏有效的键值(KV)缓存和可扩展的并行解码机制。现有的加速方法通常孤立地研究KV缓存和并行解码,忽略了在联合应用缓存重用和并行令牌验证时出现的I/O瓶颈。在本工作中,我们引入了Flash-dLLM,一个无需训练的推理加速框架,用于实现快速且内存高效的dLLMs。Flash-dLLM首先识别出GPU内存I/O是启用KV缓存的dLLM推理中的主要瓶颈,并通过一个IO感知的融合KV缓存内核来解决该问题,该内核减少了冗余的内存移动。基于这一优化的缓存机制,Flash-dLLM进一步提出了一种高效的、由KV缓存驱动的草拟与验证解码策略,其中dLLM自身同时充当草拟器和验证器,无需辅助模型。这种统一的设计使得解码速度更快,同时保持生成质量,并提高了对更长序列和更大批处理规模的扩展性。在数学推理和代码生成基准上的大量实验表明,Flash-dLLM在推理速度和内存效率方面均持续优于现有的最先进的dLLM加速方法。特别是,在GSM8K和HumanEval上,它分别比之前最强的基线Elastic-Cache实现了5.1倍和11.0倍的加速。

英文摘要

Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce $\textbf{Flash-dLLM}$, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger batch size. Extensive experiments on mathematical reasoning and code-generation benchmarks demonstrate that Flash-dLLM consistently outperforms existing state-of-the-art dLLM acceleration methods in both inference speed and memory efficiency. In particular, it achieves $5.1\times$ and $11.0\times$ speedups over prior strongest baseline Elastic-Cache on GSM8K and HumanEval, respectively.

发表机构

  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑