KV 缓存成为新的内存墙
The KV Cache Is the New Memory Wall
浏览论文内容
中文总结 AI 辅助
长上下文推理中,KV 缓存成为内存带宽瓶颈;论文提出统一分析协议,在多种硬件上评估五类 KV 优化技术,揭示三阶段结构并给出设计规则。
中文摘要 AI 辅助
长上下文下的自回归 LLM 推理受限于内存带宽,而非算术吞吐量,并且随着序列长度的增长,瓶颈资源从模型权重转移到了键值(KV)缓存。对于 BF16 格式的 Llama-3-70B,其 140 GB 的权重占用超过了单个加速器的 80 GB HBM,而一个 128k token 的序列会额外增加 42 GB 的 KV 缓存。压缩、驱逐、分页、共享或卸载 KV 状态的技术层出不穷,但报告的结果使用了不一致的工作负载、硬件和质量指标,阻碍了跨论文的比较。本 SoK 论文从分析角度统一了该领域,采用了一种严格区分派生声明和报告声明的协议。我们推导出闭式算术强度,它是上下文长度的衰减函数,并以 NVIDIA H100、NVIDIA B200 和 AMD MI300X 的硬件拓扑为参数,包括每晶粒带宽划分以及 KV 流量超过权重流量的交叉长度。我们将文献分为五个领域:量化、token 驱逐、KV 分页、前缀缓存和异构分层,并在单一协议下对每个领域的一种方法在 128k 上下文下进行评估。核心发现是一个三阶段结构:在硬件特定的交叉点以下,权重流量占主导,KV 压缩带来的加速可忽略不计;超过该交叉点后,KV 流量占主导,每个领域以质量换取接近屋顶线的带宽节省。分页和前缀共享是无损的,但解决的是容量问题,而非带宽问题。量化和驱逐直接削减带宽,其退化在低于 4 位精度时加速,并在位置敏感任务上对驱逐变得不连续。分层将带宽墙转化为受 PCIe 或 NVLink 而非 HBM 约束的互连问题。最后,我们总结了根据硬件、上下文长度和质量预算选择压缩领域的设计规则。
英文摘要
Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128k-token sequence adds 42 GB of KV cache. Techniques that compress, evict, page, share, or offload KV state have proliferated, but reported gains use inconsistent workloads, hardware, and quality metrics, preventing cross-paper comparison. This SoK paper unifies the field analytically, with a protocol that strictly separates derived and reported claims. We derive closed-form arithmetic intensity as a decaying function of context length, parameterized by hardware topology for NVIDIA H100, NVIDIA B200, and AMD MI300X, including per-die bandwidth partitioning and the crossover lengths where KV traffic overtakes weight traffic. We classify the literature into five domains, quantization, token eviction, KV paging, prefix caching, and heterogeneous tiering, evaluating one method per domain at 128k context under a single protocol. The central finding is a three-regime structure: below a hardware-specific crossover, weight traffic dominates and KV compression yields negligible speedup; beyond it, KV traffic dominates and each domain trades quality for bandwidth savings approaching the roofline bound. Paging and prefix sharing are lossless but address capacity, not bandwidth. Quantization and eviction cut bandwidth directly, with degradation that accelerates below 4-bit precision and turns discontinuous for eviction on position-sensitive tasks. Tiering converts the bandwidth wall into an interconnect problem bounded by PCIe or NVLink rather than HBM. We close with design rules for selecting a compression domain given hardware, context length, and quality budget.
发表机构
- Dell Technologies(戴尔科技集团)
机构由 AI 辅助整理,请以论文原文为准。