发表机构
Seoul National University; Stanford University(首尔大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长上下文推理中KV缓存的内存瓶颈,提出TaSQ方法,通过查询引导通道加权、跨头归一化和协方差感知通道分组定制VQ量化空间,在1比特压缩下显著优于现有基线,并支持更大批处理与更高吞吐量。
AI 中文摘要
在长上下文大语言模型推理中,键值(KV)缓存成为主要的内存瓶颈,对内存容量和带宽造成巨大压力。为缓解这一瓶颈,向量量化(VQ)已成为一种有前景的激进KV缓存压缩方法。然而,现有的VQ方法在1比特场景下性能显著下降。在如此极端的压缩下,每个码本必须用有限的质心集合表示更大的通道组,这使得有效利用其容量变得越来越具有挑战性。为解决此问题,我们引入了TaSQ,它通过结合查询引导的通道加权、跨头归一化和协方差感知的通道分组来定制VQ目标空间,以更好地反映缓存激活的误差敏感性和统计结构。由于这些变换与RoPE兼容,且可以轻松合并到投影权重和码本中,TaSQ保留了传统的VQ查找结构,并增加了可忽略的推理开销。在通用、长思维链推理和长上下文检索基准测试中,TaSQ始终优于现有的低位KV缓存VQ基线,同时保持推理稳定性。在单个RTX 6000 Ada GPU上,其SGLang实现支持高达14倍的更大批处理大小,并且与BF16基线相比,实现了1.87倍的更高峰值吞吐量。
英文摘要
The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.