arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散大语言模型的词缀缓存

Affix Cache for Diffusion Large Language Models

Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu, Jian Weng, Marco Canini

arXiv 2608.26140首次发表:更新:

发表机构

KAUST; CUHK; LUMS(阿卜杜拉国王科技大学; 香港中文大学; 拉合尔管理科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散大语言模型高效推理难题,提出面向词缀的ACache缓存机制,通过复用词缀缓存并仅重新计算约20%关键锚定词元,可降低延迟、提升吞吐量并恢复精度损失。

AI 中文摘要

扩散大语言模型(DLLMs)支持非自回归解码与双向上下文建模,但高效推理仍具挑战性。与键值(KV)缓存可复用共享前缀的自回归系统不同,DLLMs通过双向注意力将共享上下文词元的KV状态与正在生成的词元耦合,导致朴素缓存复用失效,而完全重新计算成本高昂。本文提出ACache,一种面向词缀的DLLMs共享文本跨度缓存复用机制,突破了仅适用于前缀的限制。ACache通过测量锚定词元(Anchor Tokens)对掩码生成词元的影响,识别出每个请求特有的少量关键词面子集,仅重新计算这些词元的KV状态,同时复用其余词缀缓存。基于Fast-dLLM构建的ACache,在不同设置下,当重新计算约20%的词缀词元时,可恢复直接词缀缓存复用造成的精度损失;在Nano-vLLM引擎上构建的共享前缀原型显示,ACache可将重新计算延迟降低最多55.7%,并将端到端吞吐量提升最多1.68倍。

英文摘要

Diffusion Large Language Models (DLLMs) enable non-autoregressive decoding, but efficient inference support remains immature: unlike autoregressive models, whose requests reuse a shared prefix key-value (KV) cache, DLLMs use bidirectional attention, so a shared context's KV states depend on the tokens still being decoded, leaving directly reused caches stale and full recomputation necessary. We present ACache, a cross-request cache reuse mechanism for shared spans, or affixes, at any position: prefix, infix, or suffix. ACache measures the influence of affix tokens on the masked generation region to identify a small request-specific subset as Anchor Tokens, and recomputes only their KV states while reusing the remaining affix cache. Built on state-of-the-art intra-request caching mechanisms, ACache recovers most of the accuracy lost to direct affix-cache reuse on average when recomputing around 20% of affix tokens, and at that budget preserves more accuracy than selection criteria adapted from prior cross-request cache-reuse systems. We co-design ACache with a modern inference engine, whose attention reads each request's recomputed Anchor KV states alongside one affix cache shared across concurrent requests. Against the same system with only intra-request caching, ACache cuts recompute latency by up to 56.7%, translating to as much as 1.71$\times$ end-to-end throughput, while reducing peak KV cache memory by up to 45.8%.

Comments31 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑