arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过局部分布恢复实现高精度低比特 KV 缓存量化

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration

Gradwell Dzikanyanga, Yanqi Pan, Weihao Yang, Donglei Wu, Wen Xia, Hao Huang

arXiv 2607.16248首次发表:更新:

AI 中文总结

研究长上下文大语言模型推理中 KV 缓存量化问题,提出局部分布恢复技术,实现 DGAP。该技术可检测高风险步骤并恢复 top-K 候选分布,实验证明在 Llama-3.1-8B 等模型上能提升准确率、降低分布漂移,保持低比特缓存占用。

AI 中文摘要

长上下文大语言模型推理依赖 KV 缓存避免冗余注意力计算,但会带来高内存和带宽开销。低比特 KV 缓存量化可降低成本,但严重降低质量,如一位量化使 Llama-3.1-8B 在 RULER 下准确率从 84.2%降至 47.8%。研究发现根本原因是结构化局部排序错误,即 top-K 区域 logits 分布漂移。提出局部分布恢复技术,通过检测量化 logit 特征中高局部分布风险步骤,在 token 选择前仅恢复选定 top-K 候选分布。实现 DGAP 进行局部分布恢复,有高效风险检测器和校正器。实验表明,在 Llama-3.1-8B 上,DGAP 将 K1V1 RULER 准确率从 47.8%恢复到 83.2%,将分布漂移从 0.38 降至 0.14;在多个模型中,能保持低比特 KV 缓存占用并适度增加解码开销。

英文摘要

Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit KV-cache quantization reduces this cost, yet it severely degrade quality; particularly, one-bit quantization reduces accuracy from 84.2% to 47.8% on Llama-3.1-8B under RULER. Rather than common beliefs that absolute error of logits, we find that the root cause is structured local misranking, where the distribution of logits in top-K region is drifted. We thereby propose local distribution restoration, a new technique that detects steps with high local distribution risk from quantized-logit features and restores only the selected top-K candidate distribution before token selection. We implement DGAP to achieve local distribution restoration, with efficient risk detcetors and correctors. Expeirments show that on Llama-3.1-8B, DGAP recovers K1V1 RULER accuracy from 47.8% to 83.2% and reduces distribution drift from 0.38 to 0.14; across Llama, Mistral, and Qwen models, it preserves the persistent low-bit KV-cache footprint with modest decode overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑