arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24197cs.LG

H-Spec:无草稿方KV缓存的并行推测解码

H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache

Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn, Reed Meyerson, Zhenting Qi, Tianyu Wu, Eldar Kurtic, Minlan Yu, Alexandre Marques

首次发表
浏览论文内容

中文总结 AI 辅助

H-Spec提出无草稿方KV缓存的混合Mamba-注意力并行推测解码器,通过目标上下文注入提升平均接受长度与吞吐量,降低KV缓存开销。

中文摘要 AI 辅助

推测解码通过让轻量级草稿模型预测未来令牌以供目标模型验证,无损地加速大型语言模型推理。最近的块扩散草稿模型通过并行预测多个令牌进一步降低了草稿生成延迟。然而,现有的块草稿模型将每个输入位置的目标隐藏状态投影到单独的草稿方KV缓存中,导致随并发度增长的每请求内存和KV写入开销;直接就地重用目标KV可消除此缓存,但无法在整个块内维持草稿质量。我们提出一种混合目标上下文注入方法,将直接目标KV重用与仅在最后输入位置的目标隐藏状态相结合,无需单独的草稿方KV缓存。基于此设计,我们提出H-Spec,一种混合Mamba-注意力并行草稿模型,通过互补模块消费两种目标上下文来源。Mamba模块用投影的最后令牌目标隐藏状态初始化,而注意力模块就地重用目标KV。尽管采用循环公式,Mamba的并行扫描使H-Spec保持块并行草稿生成。在三个目标模型和多种任务上,H-Spec在平均接受长度上比最佳基线提高5.0--13.3%,在批量大小为1的令牌间延迟加速上提高5.3--12.6%。在并发服务下,H-Spec在所有评估的并发级别上持续实现更高吞吐量,同时保持比基线更低的KV缓存利用率。

英文摘要

Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model. Recent block diffusion drafters further reduce drafting latency by predicting multiple tokens in parallel. However, existing block drafters project target hidden states at every input position into a separate drafter-side KV cache, incurring per-request memory and KV-write overhead that grow with concurrency; directly reusing target KVs in place removes this cache but fails to sustain draft quality throughout the block. We propose a hybrid target-context injection method that complements direct target KV reuse with target hidden states only at the last input position, requiring no separate drafter-side KV cache. Building on this design, we propose H-Spec, a hybrid Mamba-attention parallel drafter that consumes the two target-context sources through complementary modules. Mamba modules are initialized with projected last-token target hidden states, while attention modules reuse target KVs in place. Despite its recurrent formulation, Mamba's parallel scan allows H-Spec to preserve block-parallel drafting. Across three target models and diverse tasks, H-Spec improves over the best baseline by 5.0--13.3% in mean accepted length and 5.3--12.6% in batch-size-1 inter-token latency speedup. Under concurrent serving, H-Spec consistently achieves higher throughput while maintaining lower KV cache utilization than baselines across evaluated concurrency levels.

发表机构

  • Harvard University(哈佛大学)
  • Capital One(第一资本)
  • Red Hat(红帽公司)
  • ISTA(奥地利科学技术研究所)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑