arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LaCache:用于扩散大语言模型的精确缓存和精度自适应推理

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

Xingru Chen, Zelang Liang, Yongjia Ma, Jiqing Zhan, Shuling Yang, Lian Wen, Kun Zhan

arXiv 2607.16339首次发表:更新:

发表机构

Li Auto Inc.(理想汽车公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散大语言模型算子级冗余问题,提出LaCache加速框架,通过无损缓存中间结果及集成FP8量化策略,无需训练,单独使用可实现约1.3倍端到端加速,与现有方法结合可达40.2倍加速且保持任务精度。

AI 中文摘要

基于扩散的大语言模型(DLLMs)在文本生成中通过半自回归(SAR)解码实现并行生成。然而,当前方法存在严重的算子级冗余,在去噪步骤中会重新计算整个序列,忽略了前缀和掩码后缀在一个块内保持不变。我们提出了LaCache,一个无需训练的加速框架,通过无损缓存和混合精度来减轻这种冗余。具体来说,LaCache采用无损状态记忆(LSM),缓存三种类型的中间结果:用于嵌入输出的EmbedCache、用于逐令牌注意力前状态的RoPECache和用于FlashAttention内在线softmax统计的FACache。这些缓存允许模型在不变的令牌上跳过冗余计算而不改变输出。为进一步缓解内存带宽瓶颈,LaCache为FFN层集成了针对扩散过程中依赖步骤的激活分布定制的每组FP8量化策略。实验表明,单独使用LaCache比普通DLLM实现了约1.3倍的端到端加速。与现有加速方法结合时,LaCache在保持可比任务精度的同时达到了高达40.2倍的端到端加速。

英文摘要

Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequence during denoising steps, ignoring that the prefix and masked suffix remain invariant within a block. We propose LaCache, a training-free acceleration framework that alleviates this redundancy through lossless caching and mixed precision. Specifically, LaCache employs Lossless State Memoization (LSM) by caching three types of intermediate results: (i) EmbedCache for embedding outputs, (ii) RoPECache for token-wise pre-attention states, and (iii) FACache for the online softmax statistics within FlashAttention. These caches allow the model to skip redundant computation on unchanged tokens without altering the output. To further alleviate memory-bandwidth bottlenecks, LaCache inegrates a per-group FP8 quantization strategy for FFN layers, tailored to step-dependent activation distributions across the diffusion process. Experiments demonstrate that LaCache alone achieves approximately 1.3X end-to-end speedup over vanilla DLLM. When combined with existing acceleration methods, LaCache reaches up to 40.2X end-to-end speedup while maintaining comparable task accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑