arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12338cs.CLcs.LG

VFold:感知对称性的跨层值缓存压缩

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

Neha Verma, Sungwon Kim, Kenton Murray, Kevin Duh

AI总结:

本研究提出感知对称性的值缓存合并策略,可在不损害性能和增加架构开销的前提下压缩LLM的值缓存,还能与其他压缩技术结合实现更高压缩比,为内存受限场景下扩展LLM上下文窗口提供了有效方案。

AI中文摘要:

缓存键值(KV)状态可加速大语言模型(LLM)解码,但在长上下文场景下,该缓存会占据内存使用的主要部分。现有多数解决方案需对LLM架构进行修改,且会产生大量开销。本研究提出一种感知对称性的值缓存合并策略,可降低缓存内存,同时避免解码时出现有害的性能下降和架构开销。此外,研究表明该方法可与现有缓存压缩技术结合,与高比率量化或键缓存剪枝配合,达到单一方法无法实现的压缩比,且额外开销极小。最终,研究发现了值缓存中未充分利用容量的一个主要来源,为内存约束下扩展上下文窗口提供了简单且高效的方向。

英文摘要:

While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.

↑