arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WUSH-KV:基于数据自适应变换的KV缓存量化

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh

arXiv 2609.38121首次发表:更新:

发表机构

Institute of Science Technology Austria (ISTA); Red Hat AI; ETH Zürich(奥地利科学技术学院; 红帽人工智能; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WUSH-KV通过数据自适应变换对KV缓存进行低位量化,降低内存和带宽成本,在2比特下性能优于或媲美OSCAR变换。

AI 中文摘要

KV缓存的内存和带宽成本随上下文长度和批量大小增长,这限制了高效的长上下文推理。为解决这一瓶颈,我们提出了WUSH-KV,用于低位KV缓存量化。它改进了WUSH方法,该方法利用矩阵乘积中两个因子的二阶统计量构建数据感知的变换,以降低量化误差。WUSH-KV使用校准数据分别构建键和值的变换,其中值变换被折叠进模型权重中,键变换在RoPE之后应用。这些变换可以与裁剪量化器配对使用。对于其中一种量化器QuEST INT,我们证明在温和假设下,WUSH变换是近似最优的。使用该量化器时,WUSH-KV降低了逐层重构误差,并在其他测试的变换中实现了最低的端到端困惑度。在端到端评估中,我们使用OSCAR风格的百分位裁剪仿射量化将WUSH-KV集成到SGLang中。在2比特设置下,WUSH-KV在所有评估的模型和下游任务上表现与OSCAR变换相当或更优。

英文摘要

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑