AI 中文总结
该研究针对长上下文推理中KV缓存的内存瓶颈,提出注意力感知变换编码(AATC)方法,在多基准测试中实现约5.8倍压缩且保持近乎无损的准确率,优于基线方法。
AI 中文摘要
键值(KV)缓存存储来自过往token的信息,是长上下文推理中的主要内存瓶颈。现有量化方法通过用低精度数据类型统一表示KV缓存并设计量化方案以最小化缓存本身的重构误差来解决该瓶颈,但未考虑该误差如何通过注意力机制传播。我们证明,在白噪声量化模型下,期望的注意力感知失真可分解为跨token和通道的可加键与值贡献。基于信号处理和率失真理论的经典工具——变换编码与反向注水算法,我们引入注意力感知变换编码(AATC),在校准集上分配比特以最小化注意力感知失真。在Llama-3.1-8B-Instruct和Qwen-2.5-7B-Instruct模型上,于LongBench、RULER、GSM8K、MMLU-Pro和MATH-500基准测试中评估,我们的方法在约5.8倍压缩时达到近乎无损的准确率,而各基线方法在至少部分设置中出现性能下降。
英文摘要
The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the cache itself, without accounting for how that error propagates through attention mechanisms. We prove that, under a white-noise quantization model, the expected attention-aware distortion decomposes into additive key and value contributions that factor across tokens and channels. Building on transform coding and reverse water-filling, which are classical tools from signal processing and rate-distortion theory, we introduce Attention-Aware Transform Coding (AATC), which allocates bits over a calibration set to minimize attention-aware distortion. On Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, evaluated across LongBench, RULER, GSM8K, MMLU-Pro, and MATH-500, our method achieves near-lossless accuracy at approximately $5.8\times$ compression, whereas each baseline degrades in at least some settings.