arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiSparse:通过分层KV缓存管理扩展稀疏注意力解码

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management

Zhiqiang Xie, Zhangheng Huang, Tingwei Huang, Ziyi Xu, Ruiyang Ma, Christos Kozyrakis

arXiv 2608.07009首次发表:更新:

AI 中文总结

HiSparse是一种分层KV缓存,通过将KV历史存主机、GPU设固定缓存等方式,提升长上下文稀疏注意力LLM解码吞吐量达4.7倍,未增加每token成本,已集成到SGLang并在多平台验证。

AI 中文摘要

Top-k稀疏注意力使长上下文大语言模型(LLM)解码的计算成本降低:每一步仅读取数千个选中的KV条目,而非完整上下文。然而,服务系统通常会将整个KV缓存保留在GPU高带宽内存(HBM)中,以确保每个位置都可被选中,因此请求的内存开销仍会随其完整上下文长度增长——解码在计算资源耗尽前就会遭遇容量瓶颈,而KV缓存超出HBM的上下文根本无法服务。我们提出HiSparse,一种适用于稀疏注意力服务的、精确且与索引器无关的分层KV缓存。HiSparse将每个请求的完整KV历史保留在主机内存中,并通过小型固定大小的GPU缓存限制其解码内存占用;融合的CUDA内核在解码CUDA图内处理每一层的选中操作——命中检测、最近最少使用(LRU)替换以及主机到设备的取数;对于跨层共享选中项的模型,精确的分层预取可隐藏约一半的剩余未命中开销。由于仅改变KV的放置位置,模型输出保持不变。HiSparse已被合并到上游SGLang中,并在H200、B200和GH200平台上针对三类稀疏注意力家族(DSA、NSA和Quest)进行评估:它在长上下文工作负载上将峰值生成吞吐量提升了高达4.7倍,同时保持了相当的每token延迟,并降低了高负载下的首token时间——而无IO的基准测试显示,该解析机制本身未增加可测量的每token成本,因此主机-设备IO是受限驻留的唯一代价。

英文摘要

Top-k sparse attention makes long-context LLM decoding cheap to compute: each step reads only a few thousand selected KV entries rather than the full context. Serving systems, however, typically keep the entire KV cache in GPU HBM so that every position stays selectable, so a request's memory bill still grows with its full context length--decoding hits a capacity wall long before it runs out of compute, and a context whose KV cache exceeds HBM cannot be served at all. We present HiSparse, an exact, indexer-agnostic hierarchical KV cache for sparse-attention serving. HiSparse keeps each request's full KV history in host memory and bounds its decode footprint with a small, fixed-size GPU cache; a fused CUDA kernel resolves each layer's selections--hit detection, LRU replacement, and host-to-device fetches--inside the decode CUDA graph; and, for models that share selections across layers, exact layer-wise prefetching hides roughly half of the remaining miss overhead. Because only KV placement changes, model outputs are unchanged. HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high load--and a no-IO oracle shows the resolution mechanism itself adds no measurable per-token cost, leaving host-device IO as the only price of bounded residency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑