D-Quant:用于KV缓存量化的可漂移熵编码
D-Quant: Driftable Entropy Coding for KV Cache Quantization
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对KV缓存量化中固定宽度编码信息损失严重的问题,提出D-Quant框架,通过漂移机制将熵编码转为固定大小比特流,实现高效并行反量化,兼顾压缩率与部署性能。
AI中文摘要:
KV缓存已成为部署大型语言模型的主要瓶颈,因为其内存占用随序列长度和批处理大小线性增长,对内存容量和带宽施加了巨大压力。在各种KV缓存压缩技术中,量化因其有效性和易于部署而特别有吸引力。然而,大多数现有方法依赖于固定宽度量化,其中b位表示固有地限于2^b个量化级别。随着位宽减小,可用级别的数量呈指数级收缩,导致严重的信息损失和快速性能下降。我们进一步观察到,固定宽度量化未能利用KV缓存的高度非均匀分布。在旋转和归一化之后,KV值近似遵循正态分布,大多数值集中在中心附近,只有一小部分出现在尾部。然而,固定宽度编码为频繁和稀有符号分配相同数量的位。熵编码通过为频繁符号分配较短的码字和为稀有符号分配较长的码字,自然地利用了这种非均匀性,显著减少了表示所需的平均位数。然而,其可变长度输出不适合高度并行的注意力内核,其中有效的反量化和计算依赖于规则的内存布局和固定步长访问。为了弥合这一差距,我们提出了D-Quant,一个灵活的KV缓存量化框架,引入了漂移机制,将每个令牌的熵编码表示转换为固定大小的比特流,从而在注意力内核内实现规则的内存访问和并行反量化。
英文摘要:
The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness and ease of deployment. However, most existing methods rely on fixed-width quantization, where a $b$ bit representation is inherently limited to $2^b$ quantization levels. As the bit width decreases, the number of available levels shrinks exponentially, leading to severe information loss and rapid performance degradation. We further observe that fixed-width quantization fails to exploit the highly non-uniform distribution of KV cache. After rotation and normalization, KV values approximately follow a normal distribution, with most values concentrated near the center and only a small fraction appearing in the tails. Nevertheless, fixed-width coding allocates the same number of bits to frequent and rare symbols. Entropy coding naturally exploits such non-uniformity by assigning shorter codewords to frequent symbols and longer ones to rare symbols, substantially reducing the average number of bits required for representation. However, its variable-length output is not suited to highly parallel attention kernels, where efficient dequantization and computation rely on regular memory layouts and fixed-stride accesses. To bridge this gap, we propose \textbf{D-Quant}, a flexible KV cache quantization framework that introduces a \textbf{drift} mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.