发表机构
KAIST; GIST; Google Research(韩国科学技术院; 光州科学技术院; 谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Dual-QK方法,通过非正交查询和键变换解决旋转量化中查询能量分散问题,实现2位KV缓存压缩与通道剪枝,在长上下文任务中优于OSCAR,显著提升吞吐量。
AI 中文摘要
长输入和扩展生成增加了键值(KV)缓存的存储和访问成本。低位量化减少了存储和内存流量,而查询通道剪枝可以进一步减少键缓存的读取。基于旋转的量化将键离群值的能量重新分布到各通道。为了保持计算不变性,必须对查询应用相同的正交变换,以保持查询-键点积。然而,这种旋转可能分散查询能量,削弱保留少数大分量与剪除许多小分量之间的区分度。我们提出了Dual-QK,它使用成对的非正交查询和键变换来解决这一冲突。利用校准的查询和键统计信息,Dual-QK将部分键白化与查询对齐基相结合,以平衡INT2量化的键尺度并集中查询能量用于动态通道剪枝。通道0保护和桶相对RoPE支持长上下文中的低位精度。在四个模型、五个生成基准和长上下文检索任务上的实验表明,在40%查询通道稀疏度下,大多数任务上的准确率优于OSCAR。在128K上下文中,Dual-QK相对于未剪枝的BF16提供了6.8倍的KV缓存压缩和估计8.3倍的KV读取量减少。在评估配置下,我们的SGLang实现实现了相对于未剪枝BF16高达3.75倍的解码吞吐量。
英文摘要
Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to $3.75\times$ the decoding throughput of unpruned BF16.