发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出名为NOVA-KV的KV缓存量化方法,将其建模为变换编码问题,推导闭式最优变换,在每元素2比特时可恢复标量量化方法损失的大部分长上下文检索精度。
AI 中文摘要
长上下文大语言模型(LLM)的解码过程每一步都会读取键值(KV)缓存,加载该缓存的时间长于对其计算注意力的时间,因此吞吐量受带宽限制。由此,减小缓存尺寸可同时提升解码速度与服务容量。挑战在于需在减小缓存尺寸的同时保留注意力乘积、保持重构成本低廉,并使用固定的每token比特数。在每元素2比特的设置下,最具竞争力的方法依赖正交变换,但现有技术要么对数据不敏感,要么使用查询统计量却未从失真准则推导变换,且这些技术依赖基于随机或哈达玛旋转构建的变换,此类变换会均衡各维度的方差而非压缩能量,还使用低速率下次优的固定宽度标量量化器。本文将KV缓存量化建模为变换编码问题,其中失真为注意力乘积的误差;在高分辨率模型下,从校准统计量推导键和值的闭式最优变换,发现最优键变换非正交且满足广义帕塞瓦尔关系:注意力感知失真变为变换域的均方误差(MSE),因此可将MSE最优向量量化器直接应用于变换后的键系数。为满足固定宽度布局要求,证明在相同高分辨率模型下,将系数分组为等体积划分可使等大小码本达到可变速率最优值。在每元素2比特时,本文提出的名为NOVA-KV的方法可恢复标量量化方法在相当吞吐量下损失的大部分长上下文检索精度。
英文摘要
Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduce cache size while preserving the attention products, keeping reconstruction cheap, and using a fixed per-token bit count. At two bits per element, the most competitive methods rely on orthogonal transforms. However, existing techniques are either data-oblivious or use the query statistics without deriving the transform from a distortion criterion. Moreover, they rely on transforms built on top of random or Hadamard rotations, which equalize variances across entries rather than compacting energy, and fixed-width scalar quantizers, which are suboptimal at low rates. In this paper, we formulate KV cache quantization as a transform coding problem in which distortion is the error in the attention products. We derive closed-form optimal transforms for keys and values from calibration statistics, under a high-resolution model. We show that the optimal key transform is not orthogonal and satisfies a generalized Parseval relation: the attention-aware distortion becomes mean-squared error (MSE) in the transform domain. Thus, we can use MSE-optimal vector quantizers applied directly to the transformed key coefficients. To meet the fixed-width layout requirement, we show that grouping coefficients into equal-volume partitions makes equal-size codebooks attain the variable-rate optimum under the same high-resolution model. At two bits per element, our method, termed NOVA-KV, recovers most of the long-context retrieval accuracy lost by scalar quantization methods at comparable throughput.