arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KV-COBRA:通过协同优化的位-秩分配实现KV缓存压缩

KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon

arXiv 2609.24298首次发表:更新:

发表机构

Pohang University of Science and Technology (POSTECH)(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KV-COBRA通过按注意力头协同优化秩和位宽分配,在低比特率下实现更优的KV缓存压缩,显著减少精度损失且无额外开销。

AI 中文摘要

是什么限制了极端比特率下的KV缓存压缩?我们认为,限制因素并非压缩方案的选择,而是其预算在注意力头之间的分配方式。现有方法统一应用秩和位宽,忽略了每个注意力头在秩截断与量化之间存在不同的最优组合。我们证明,仅使用标准的低秩投影和标量量化,对每个注意力头协同优化秩和位宽,即可优于统一分配方案,且在低比特率下收益最大。我们的方法KV-COBRA(协同优化的位-秩分配)将此形式化为一个资源分配问题:它在每个注意力头内平衡秩截断损失与量化损失,然后在各注意力头之间重新分配预算以最小化总失真。融合的Hadamard旋转均衡了各通道的方差,并按注意力KL重要性对SVD基进行重排序,使求解器具有查询感知能力。同一分配器可扩展至联合的K+V压缩。在困惑度、零样本和长上下文基准上,从每维0.5到4比特(bpd),KV-COBRA在低bpd下显示出评估方法中最小的精度下降,且无每令牌开销。

英文摘要

What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head, using only standard low-rank projection and scalar quantization, dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint $K{+}V$ compression. On perplexity, zero-shot, and long-context benchmarks from $0.5$ to $4$ bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑