用于KV缓存的JoLT:通过联合塔克分解和JL残差分配实现大语言模型的近无损KV缓存压缩
A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation
查看机构详情
- Universität Trier(特里尔大学)
- Fachbereich IV, Mathematik, Universität Trier(特里尔大学数学四系)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对KV缓存内存成本高的问题,提出JoLT方法,通过联合塔克分解和JL残差分配压缩缓存,单个拉格朗日对偶分配参数,实现近无损2 - 3倍压缩,随机变体FlashJoLT还加速了压缩时间。
中文摘要 AI 辅助
键值(KV)缓存已成为Transformer推理中主要的内存成本。它随批量大小、上下文长度和深度增长,在长上下文时,它而非模型权重限制了吞吐量。有两类方法可降低其成本。低秩方法对缓存的二维切片进行分解,量化方法降低每个条目的位宽。但两类方法都未利用层缓存是三阶张量这一事实,其三个轴(头、令牌和特征)具有不同程度的冗余。本文直接采用张量视角。我们的方法JoLT应用部分塔克分解,仅压缩令牌和特征轴,同时保留头和层轴不变,然后用约翰逊 - 林登施特劳斯(JL)旋转低比特残差恢复截断丢弃的能量。单个拉格朗日对偶在一个字节预算下,按层组并分别针对键和值分配塔克秩和残差位宽。结果是实现了近无损的2 - 3倍压缩:在分组查询注意力模型(Mistral - 7B - v0.3)和多头注意力模型(LLaMA - 2 - 13B)上,困惑度、GSM8K准确率和RULER针在草堆检索结果都保持在未压缩基线的统计噪声范围内或与之相当。在2倍压缩时,JoLT在两种架构上重建缓存的相对弗罗贝尼乌斯误差为0.009(键)和0.006(值),大致比跨层奇异值分解和4位量化低一个数量级。随机奇异值分解变体FlashJoLT在匹配质量下实现了5 - 13倍的压缩时间加速。
英文摘要
The key-value (KV) cache is the dominant memory bottleneck in long-context language model inference. Existing compression methods apply low-rank factorization or quantization independently, without jointly allocating rank and precision under a shared storage budget. We introduce JoLT, a training-free compressor that treats grouped prefill caches as fourth-order tensors and applies partial Tucker decomposition along the token and feature modes, the two axes that carry low-rank structure, while leaving the head and layer modes intact. A rotated low-bit quantizer captures the truncation residual, and a single Lagrangian dual allocates per-group Tucker ranks and residual bit-widths under a global byte constraint. FlashJoLT replaces the exact token-mode SVD with a randomized approximation that matches JoLT within the free zone at a fraction of the compression cost, and a fused Triton decode kernel evaluates attention directly over the stored factors without materializing dense KV tensors. Across five models from four architecture families, covering multi-head attention, grouped-query attention, and mixture-of-experts architecture, JoLT achieves 2 - 3x compression with less than 0.2% perplexity degradation, without retraining. On RULER at 64K context with LLaMA-3.1-8B, retrieval accuracy remains near-lossless through 3x and declines by only 0.90 and 2.40pp at 4x and 5x, respectively. JoLT demonstrates that tensor-aware low-rank decomposition and quantized residuals, unified under a single storage budget, achieve near-lossless KV-cache compression across diverse model architectures without retraining.