arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强草稿需要紧凑记忆:带压缩KV缓存的长上下文推测解码

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

Tong Yuan, Chengxi Liao, Zeyi Wen

arXiv 2608.30252首次发表:更新:

发表机构

Data Science and Analytics Thrust, Information Hub; The Hong Kong University of Science and Technology (Guangzhou)(数据科学与分析方向、信息枢纽; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出记忆增强草稿技术,为长上下文推测解码配备压缩KV记忆,可降低70%以上草稿侧内存,使Llama~3.1-8B和70B模型分别实现最高2.08倍、3.33倍解码加速,同时保持推测解码的无损保证。

AI 中文摘要

文档摘要和多轮智能体等长上下文大语言模型(LLM)应用需要对数万token的前缀进行生成,解码延迟成为主要瓶颈。推测解码(SD)在不改变模型输出的情况下降低延迟,但其加速效果取决于被接受的草稿token数和草稿步骤延迟:轻量草稿速度快但缺乏捕捉长程依赖的能力,而强独立草稿虽能提升接受率,但在长前缀下会带来不断增长的KV访问成本。本文提出面向长上下文SD的记忆增强草稿技术,为强独立草稿配备压缩的草稿侧KV记忆:轻量适配器构建并增量更新该记忆,以保留远程信息和精确的近期上下文;目标验证器保留其完整KV缓存并采用标准接受/拒绝规则,从而保持SD的无损保证。在Llama~3.1-8B和70B目标模型上,针对最长32K的前缀长度开展的实验表明,本文方法可将草稿侧内存降低70%以上,分别比自回归解码实现最高2.08倍和3.33倍的加速。

英文摘要

Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a major bottleneck. Speculative decoding (SD) reduces latency without changing model outputs, but its speedup depends on both accepted draft tokens and draft-step latency: Lightweight drafts are fast but lack the capacity to capture long-range dependencies, whereas strong independent drafts recover acceptance but incur growing KV-access cost at long prefixes. We introduce memory-augmented drafting for long-context SD, equipping a strong independent draft with compressed draft-side KV memory: A lightweight adaptor constructs and incrementally updates this memory to retain distant information and exact recent context. The target verifier retains its full KV cache and applies the standard accept/reject rule, preserving SD's lossless guarantee. Experiments on Llama~3.1-8B and 70B targets at prefix lengths up to 32K show that our method reduces draft-side memory by over 70%. It achieves speedups of up to 2.08x and 3.33x , respectively, over autoregressive decoding.

CommentsEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑