发表机构
Nokia; Rice University(诺基亚; 莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPECTRA是一种无需训练的KV缓存压缩编解码器,通过谱变换编码将比特预算集中于关键通道,在Llama-3.1-8B等模型上实现最高12倍压缩,突破2比特极限,提升长上下文处理能力
AI 中文摘要
在智能体时代,大型语言模型(LLMs)越来越多地处理长输入,从整篇文档、代码库到多轮对话。它们的推理内存由键值(KV)缓存主导,即模型读取和生成的每个token的注意力键与值的存储。由于缓存随上下文长度增长,且每生成一个token都要完整重读,更长的上下文意味着更多GPU内存。为降低该成本,现有多数方法通过将每个存储值降低到相同低精度来压缩KV缓存,该技术称为量化。这些方法可将压缩推至每个值近2比特,但很少能进一步,因为在这一2比特极限处质量骤降:四个层级对于缓存的异常值密集值来说过少,少数大条目占用层级,其余则坍缩为噪声。自然的补救方法是在重要通道(特征维度)上分配更多比特,其余则分配更少,但原始缓存无此控制:其通道相关性极强,没有哪个通道更突出。我们的分析表明,当缓存旋转至由其自身统计量计算的坐标系时,该控制会出现,消除了这些相关性。在此坐标系中,一小部分通道承载几乎所有信息,将比特预算分配给这些少数通道比均匀分配准确得多。基于此分析,我们开发了SPECTRA,一种无需训练、可直接插入的编解码器,它将缓存重新编码至该坐标系,并将比特预算集中在承载信号的通道上。在Llama-3.1-8B和Qwen2.5-7B的长上下文基准测试中,SPECTRA在4倍压缩时近乎无损,在8倍压缩时表现与均匀量化相当(此时均匀量化已失效),且可达12倍压缩,将可用压缩率推过2比特极限,使同一GPU可容纳更长上下文和更大批次。
英文摘要
Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Their inference memory is then dominated by the key-value (KV) cache, the stored attention keys and values of every token the model has read and generated. Because the cache grows with context length and is re-read in full at every generated token, a longer context means more GPU memory. To reduce this cost, most existing methods compress the KV cache by lowering every stored value to the same low precision, a technique known as quantization. They can push this to nearly two bits per value, but rarely further, because quality drops sharply at this 2-bit cliff: four levels are too few for the cache's outlier-heavy values, where a few large entries consume the levels and collapse the rest into noise. A natural remedy is to spend more bits on the channels (feature dimensions) that matter and fewer on the rest, but the raw cache offers no handle: its channels are strongly correlated, so none stands out as more important. Our analysis shows that this handle appears once the cache is rotated into a coordinate system computed from its own statistics, removing these correlations. There, a small fraction of channels carries almost all the information, and spending the budget on those few is far more accurate than spreading it evenly. Guided by this analysis, we develop SPECTRA, a training-free, drop-in codec that re-encodes the cache into this coordinate system and concentrates the bit budget on the channels that carry the signal. On Llama-3.1-8B and Qwen2.5-7B over long-context benchmarks, SPECTRA is near-lossless at 4x compression, competitive at 8x where uniform quantization has collapsed, and reaches up to 12x, pushing usable compression past the 2-bit cliff so the same GPU holds longer contexts and larger batches.
Comments28 pages