arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

窗口化多令牌预测:消除百万令牌上下文下的全上下文草稿键值开销

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Alagappan Valliappan

arXiv 2607.21535首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究百万令牌上下文下MTP草稿头成本过高问题,提出窗口化MTP方法,仅对草稿注意力应用滑动窗口等,可降低解码成本、改善延迟,在多种架构上效果显著,还能回收未读取草稿键值。

AI 中文摘要

推测性解码通过让低成本的草稿提议令牌并由目标并行验证来加速自回归生成。前沿模型越来越多地内置多令牌预测(MTP/NEXTN)草稿头,假设草稿成本可忽略不计。但在百万令牌上下文时此假设不成立,MTP草稿头在每个草稿步骤通常对整个键值缓存进行全注意力计算,其读取成本随上下文线性增长并主导草稿成本。我们仅对草稿的注意力应用流式语言模型风格的滑动窗口加注意力汇聚(窗口化MTP),保持全注意力验证不变。它无需训练、即插即用且无损,将草稿的键值工作集限制为常数,在100万个令牌时减少约99%的键值条目。在单GPU上的SGLang中,针对三种架构系列在100万个令牌上下文下,窗口化使每个解码步骤的成本比原生MTP草稿降低28%至44%,端到端解码延迟也相应改善,同时保留目标的验证输出分布,且未读取的草稿键值可通过紧凑环形缓冲区回收而不影响接受率或质量。

英文摘要

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with context and comes to dominate the draft cost -- precisely where speculation is most valuable. The effect compounds with draft length (a deep native draft can turn net-negative, slower than no speculation) and sharpens under hybrid/linear-attention targets, where cheaper verification leaves the draft's full-attention read exposed. We apply a StreamingLLM-style sliding window plus attention sink to the draft's attention only (Windowed-MTP), leaving full-attention verification intact. It is training-free, drop-in, and lossless by construction: the full-attention target still decides every accepted token, so windowing changes only which tokens are proposed, never which are accepted. It bounds the draft's KV working set to a constant, dropping ~99% of KV entries at 1M. Across three architecture families (Qwen GDN-MoE 35B/122B and a Mamba2-hybrid NoPE 120B) at 1M context on a single GPU in SGLang, windowing cuts the per-decode-step cost over the shipping native MTP draft by +28% to +44%, an input-invariant margin that widens with context. Since per-token latency is this cost divided by acceptance length, at matched acceptance end-to-end decode latency improves by the same amount, and more where windowing also lifts acceptance, while preserving the target's verified output distribution. Finally, the unread draft KV -- 7.7-11% of total KV at 1M -- is reclaimed via a compact ring buffer at no acceptance or quality cost.

Comments25 pages, 2 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑