arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02815cs.AI

iS-KV:基于块增量SVD的在线低秩KV缓存压缩

iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD

Yiren Zhao, Guanghui Song, Tianrui Qin, Kejiang Ye, Cheng-zhong Xu, Xitong Gao

首次发表
浏览论文内容

中文总结 AI 辅助

iS-KV通过块增量SVD实现在线低秩KV缓存压缩,保持近期窗口精确并同步历史坐标,在多个模型上显著压缩缓存且准确率接近原始模型,优于token驱逐方法。

中文摘要 AI 辅助

长链思维推理在自回归解码过程中显著增加了KV缓存的内存占用,因为每个生成的token都会引入新的键和值状态,导致缓存随解码长度线性增长。现有的KV缓存压缩方法通常通过token驱逐来控制这种增长,但不可逆的删除可能会移除后续推理可能需要重新访问的历史状态。基于SVD的低秩压缩提供了一种替代方案,通过更紧凑的表示保留所有位置。然而,将其从固定提示缓存扩展到在线解码并非易事。通过我们的研究,我们发现如果为新token更新基,而旧token保留在旧基中的坐标,存储的历史会显著漂移。基于这一观察,我们提出了iS-KV,一种用于长时程推理的在线低秩KV缓存压缩方法。iS-KV保持一个近期窗口精确,同时将较旧的状态增量折叠为有界秩表示。随着低秩基的演化,它同步历史坐标与更新后的基以保持表示一致性。在DeepSeek-R1-Distill-Llama-8B上,iS-KV在4.06倍持久KV压缩下实现了82.6%的准确率,接近原始模型的83.6%。在Qwen3-8B上,它在5.64倍压缩下实现了89.2%的准确率。在匹配的内存预算下,iS-KV始终优于token驱逐基线。

英文摘要

Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model's 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Guangdong OPPO Mobile Telecommunications Corp., Ltd.(广东欧珀移动通信有限公司)
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • University of Macau(澳门大学)
  • Shenzhen University of Advanced Technology(深圳先进技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑