AI 中文总结
PuzzleKV将KV缓存划分为固定长度逻辑页,以页为独立压缩单元,在匹配存储预算下实现高压缩率与近似完整KV的性能,在RULER、LongBench等基准上表现优于Global SVD。
AI 中文摘要
大型语言模型(LLMs)的长上下文推理日益受到键值(KV)缓存所需内存的限制。KV缓存压缩通过降低先前token的存储成本解决该问题,在现有方法中,低秩压缩极具吸引力,因为它能以降维形式表示每个token。以往的低秩方法通常从模型权重中推导固定投影空间、从校准激活中构建固定空间,或在宽泛的缓存区域上构建共享基,此类表示可能无法捕捉重要的细节信息。我们将每个头的KV缓存划分为固定长度的逻辑页,并观察到单个页面内存在显著的低秩结构。基于该观察,我们提出PuzzleKV,这是一种无需训练和校准的方法,将每个完成的页面视为独立的压缩单元。PuzzleKV在每层和每个KV头内分解页面,直接在稠密且分解后的页面上计算注意力,并在自回归解码期间增量压缩新的合格页面。跨模型、上下文长度和基准的实验表明,在匹配的存储预算下,PuzzleKV具有有效性。在约为原始KV缓存60%的存储量下,PuzzleKV在所有评估模型和所有基准设置中达到了完整KV性能的96%以上,在RULER上相较于Global SVD有显著提升,在LongBench上表现具有竞争力。为实现更激进的压缩率,PuzzleKV可进一步与量化结合,仅使用原始存储的18.7%即可保留完整KV性能的93%以上。
英文摘要
Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches, low-rank compression is particularly attractive because it represents every token in reduced dimensions. Previous low-rank methods typically derive fixed projection spaces from model weights, construct fixed spaces from calibration activations, or construct a shared basis over a broad cache region. Such representations may not capture detailed but important information. We partition each per-head KV cache into fixed-length logical pages and observe substantial low-rank structure within individual pages. Based on this observation, we propose PuzzleKV, a training- and calibration-free method that treats each completed page as an independent compression unit. PuzzleKV decomposes pages within each layer and KV head, computes attention directly over dense and factorized pages, and incrementally compresses newly eligible pages during autoregressive decoding. Experiments across models, context lengths, and benchmarks demonstrate the effectiveness of PuzzleKV under matched storage budgets. At approximately 60% of the original KV cache storage, PuzzleKV achieves more than 96% of Full KV performance across both evaluated models and all benchmark settings, with substantial gains over Global SVD on RULER and competitive performance on LongBench. To achieve a more aggressive compression ratio, PuzzleKV can be further combined with quantization while retaining more than 93% of Full KV performance using only 18.7% of the original storage.