arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33184cs.AR

面向长上下文LLM服务的高资源效率投机解码

Resource-Efficient Speculative Decoding for Long-Context LLM Serving

Fei Li, Song Liu, Shiqiang Nie, Jinyu Wang, Weiguo Wu

首次发表
浏览论文内容

中文总结 AI 辅助

SpecStream通过提前开始验证、利用KV传输计算气泡并发草稿及目标优先调度,在长上下文投机解码中提升吞吐量,分别达1.41倍和1.32倍,每GPU输出吞吐量提升55.4%。

中文摘要 AI 辅助

投机解码通过并行验证来自草稿模型的多个令牌,减少了顺序目标模型调用。然而,在受限的GPU内存下,KV缓存的增长限制了长上下文服务。将KV卸载到CPU内存缓解了这一压力。然而,现有的卸载方案在注意力计算前恢复完整的KV历史,未能充分利用验证轮内跨查询的KV共享优势。现有的并行投机解码方法也忽视了在内存带宽受限的目标验证等待历史KV时闲置的GPU计算能力。我们提出SpecStream,一种投机解码系统,它在无需等待完整KV历史恢复的情况下开始验证,并利用KV传输期间的计算气泡在同一GPU上进行并发草稿生成。它仅卸载目标提交的历史,将候选回滚保留在GPU本地。每个流式KV块服务于轮内所有查询,通过在线softmax保持完整注意力。目标优先调度在资源限制下控制草稿并发执行,以限制对目标验证的干扰。实验表明,SpecStream在有限GPU内存下支持更多并发请求的同时,保持了接近SGLang投机解码的任务质量。在不同数据集上,与卸载基线相比,Qwen3和InternLM2.5的平均吞吐量分别提升了1.41倍和1.32倍。与在分离的目标和草稿GPU上进行并行投机解码相比,SpecStream每GPU的输出吞吐量平均提升了55.4%。

英文摘要

Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. Offloading KV to CPU memory relieves this pressure. However, existing offloading schemes restore the full KV history before attention and fail to fully exploit the benefits of KV sharing across queries within a verification round. Existing parallel speculative decoding methods also overlook idle GPU compute capacity while memory-bandwidth-bound Target verification waits for historical KV. We present SpecStream, a speculative decoding system that begins verification without waiting for the full KV history to be restored and exploits compute bubbles during KV transfers for concurrent drafting on the same GPUs. It offloads only Target-committed history, keeping candidate rollback local to the GPUs. Each streamed KV chunk serves all queries in the round, with online softmax preserving full attention. Target-priority scheduling controls Draft concurrent execution under resource limits to constrain interference with Target verification. Experiments show that SpecStream maintains task quality close to SGLang speculative decoding while supporting more concurrent requests under limited GPU memory. Across different datasets, it achieves average throughput speedups of 1.41$\times$ and 1.32$\times$ over the offloading baseline for Qwen3 and InternLM2.5, respectively. Compared with parallel speculative decoding on separate Target and Draft GPUs, SpecStream improves output throughput per GPU by an average of 55.4%.

发表机构

  • School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑