发表机构
National School of Artificial Intelligence (ENSIA)(国家人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型语言模型KV缓存内存占用瓶颈问题,提出无训练的SelKV双组件框架,通过软余弦门和注意力比率补偿机制进行缓存压缩,在LongBench数据集上仅保留25%缓存时性能强大,在多文档问答任务中表现出色且解码加速。
AI 中文摘要
大型语言模型(LLMs)自回归生成文本时依赖键值(KV)缓存,其内存占用随上下文长度线性增长,成为主要瓶颈。近期压缩方法通过令牌合并减轻成本,但常采用无差别聚合,导致表征退化和注意力偏移。本文提出无训练的双组件框架用于KV缓存压缩。一是软余弦门基于值向量相似度自适应调制合并决策,抑制或丢弃不相似令牌以保持语义保真度;二是引入注意力比率补偿机制,利用预填充注意力统计得出解码时的对数几率偏差,纠正合并引起的softmax不平衡。在LongBench(16个英语数据集)上评估,仅保留25%的KV缓存时,该框架相对于代表性的一次性基线实现了强大的压缩性能,在分组查询注意力(GQA)模型上尤其稳健,保持了近乎无损的生成质量。此外,该方法在复杂多文档问答任务上优于全缓存基线,并在100k令牌时实现了3.3倍的解码加速。
英文摘要
Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major bottleneck. Recent compression methods mitigate this cost via token merging; however, these approaches often rely on indiscriminate aggregation, which degrades representations and introduces attention sag, a mismatch where merged tokens receive the same softmax mass as individual tokens despite encoding multiple inputs. We propose a training-free, dual-component framework for KV cache compression that addresses these limitations. First, a soft cosine gate adaptively modulates merging decisions based on value-vector similarity, suppressing or discarding dissimilar tokens to preserve semantic fidelity. Second, we introduce an attention-ratio compensation mechanism that applies a decoding-time logit bias derived from prefill attention statistics, correcting the softmax imbalance induced by merging. Evaluated on LongBench (16 English datasets) while retaining only 25% of the KV cache, our framework achieves strong compressed performance against representative one-shot baselines. It is especially robust on the evaluated grouped-query attention (GQA) models, maintaining nearlossless generation quality. Furthermore, the method outperforms the full-cache baseline on complex multi-document QA tasks and delivers a 3.3x decoding speedup at 100k tokens.