arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02953cs.LG

SlimKV:基于无重建信标注意力的联合令牌-特征KV缓存压缩

SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention

  • University of Science and Technology of China(中国科学技术大学)
  • Nanyang Technological University(南洋理工大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu

AI总结:

针对长上下文LLM的KV缓存内存瓶颈,提出联合令牌-特征压缩方法SlimKV,利用低秩感知训练和位置不对称性实现无重建注意力,在保持高压缩比下性能领先并显著加速解码。

AI中文摘要:

长上下文LLM服务日益受到KV缓存内存的瓶颈限制,尤其是在资源受限的场景中。在现有的KV缓存压缩策略中,令牌级方法减少了缓存状态,但通过驱逐或冷凝存在信息丢失的风险,而特征级方法减少了每个令牌的KV维度,但可能需要全维度重建以应用位置嵌入,限制了解码加速。我们引入了SlimKV,一种与问题无关的联合令牌-特征KV缓存压缩方法。SlimKV使用低秩感知训练将长上下文压缩为具有潜在KV表示的信标记忆状态,并结合层自适应秩分配。我们进一步揭示了位置不对称性:移除键侧RoPE对信标和原始令牌的影响不同,对信标令牌的退化要小得多。利用这种不对称性,SlimKV在无K-RoPE约束下训练信标KV投影,并在解码期间实现潜在空间注意力,减轻重建延迟。在LongBench上,SlimKV在16倍/32倍压缩下优于基线,并在4倍/8倍压缩下保持领先,保留了未压缩模型超过96%的分数。大海捞针测试确认了跨证据位置的鲁棒性,效率评估显示在128K长度下,与未压缩模型相比,注意力加速高达7.34倍,端到端解码加速高达3.38倍。

英文摘要:

Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduce per-token KV dimensions but can require full-dimensional reconstruction to apply positional embedding, limiting decoding speedups. We introduce SlimKV, a question-agnostic joint token-feature KV-cache compression method. SlimKV uses low-rank-aware training to compress long contexts into beacon memory states with latent KV representations, together with layer-adaptive rank allocation. We further uncover a positional asymmetry: removing key-side RoPE affects beacon and raw tokens differently, with much smaller degradation for beacon tokens. Exploiting this asymmetry, SlimKV trains beacon KV projections under a K-RoPE-free constraint and enables latent-space attention during decoding, mitigating reconstruction latency. On LongBench, SlimKV outperforms baselines at 16x/32x compression and remains leading at 4x/8x, where it retains over 96% of the uncompressed model's score. Needle-in-a-Haystack confirms robustness across evidence positions, and efficiency evaluation shows up to 7.34x attention speedup and 3.38x end-to-end decoding speedup over the uncompressed model at 128K length.

补充信息

↑