arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24331cs.LG

DynaCalKV:通过头部分组和自适应秩分配实现键值缓存压缩

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

Tan T. Nguyen, Quan V. Dang

AI总结:

针对大语言模型KV缓存瓶颈问题,提出DynaCalKV框架,通过基于CKA相似性动态分组注意力头并自适应分配秩预算处理键缓存,值缓存采用ReCalKV方法优化,实验表明能减少键缓存参数且保持准确率,对MHA模型效果好,对GQA模型应用需更谨慎。

AI中文摘要:

随着大语言模型推理阶段需要处理长上下文窗口,键值(KV)缓存起初似乎能应对挑战,但随着上下文窗口持续增长最终成为显著瓶颈。低秩压缩是减少KV缓存内存并维持模型性能的有效方法。然而,现有方法中只有少数能区别对待键缓存和值缓存,且多采用固定注意力头分组,未充分利用注意力头间的结构相似性。本文提出改进的低秩KV缓存压缩框架。对于键缓存,基于中心核对齐(CKA)相似性动态分组注意力头,并在参数预算下自适应分配秩预算;对于值缓存,采用与ReCalKV相同方法,通过离线校准优化低秩分解以提高重构质量。在三个指令微调的大语言模型上的实验结果表明,该方法在保持竞争力的准确率的同时减少了键缓存参数数量。还观察到该策略对多头注意力(MHA)模型特别有效,而对分组查询注意力(GQA)模型应用时应更保守,尤其是在长上下文设置中。

英文摘要:

As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow. Low-rank compression has recently been studied as an effective approach to reduce KV cache memory while maintaining model performance. However, only a few existing methods treat the Key and Value caches differently, despite their distinct roles. Moreover, these methods typically employ fixed attention-head grouping, which may not fully exploit the structural similarity among attention heads. In this paper, we propose an improved low-rank KV cache compression framework. For the Key cache, we dynamically group attention heads based on Centered Kernel Alignment (CKA) similarity and allocate the rank budget adaptively under a parameter budget. For the Value cache, we adopt the same approach as ReCalKV, refining the low-rank decomposition through offline calibration to improve reconstruction quality. Experimental results on three instruction-tuned LLMs show that our method reduces the number of Key cache parameters while maintaining competitive accuracy. We further observe that the proposed strategy is particularly effective for Multi-Head Attention (MHA) models, whereas it should be applied more conservatively to Grouped-Query Attention (GQA) models, especially in long-context settings.

↑