arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28911cs.LGcs.CLcs.ITmath.IT

SemKV:面向长上下文大语言模型推理的、由质量悬崖引导的语义混合精度KV缓存量化

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

Daeha Lee, Do-Hyung Kim, Jae-Hong Kim

首次发表
浏览论文内容

中文总结 AI 辅助

SemKV基于质量悬崖提出语义混合精度KV缓存量化,保留所有token并分配特定精度,实现6.0倍存储压缩且无质量损失,优于FP16 token剪枝,替换量化器可提升压缩比至7.9倍。

中文摘要 AI 辅助

键值(KV)缓存是长上下文大语言模型(LLM)推理的主要内存瓶颈,其内存占用随上下文长度线性增长。我们表明,分数比特网格上的均匀KV量化不会平滑降级:在预先指定的多种子统计协议下,采用仿射量化器的Llama-3.1-8B-Instruct,在每个值的代码比特低至2.322时,其性能与FP16 KV在统计上无法区分,而在2.0比特时发生崩溃——这一质量悬崖出现在(2.0, 2.322]区间内,且会在生成时量化、多轮对话中重现,并可迁移至Mistral-7B。该悬崖重新定义了重要性感知混合精度:在悬崖以上,八个模型内部重要性指标在统计上可互换,因此混合精度的优势在于网格插值,可实现均匀量化无法达到的平均精度。SemKV保留所有token,按模型内部评分对token排序,并分配两个相邻的高于悬崖的精度,实现了实测6.0倍的存储减少,且与完整KV(n=900,三个种子)在统计上无检测到的质量差异,在内存预算大1.5倍的情况下,优于FP16 token剪枝。将仿射基础替换为失真优化量化器(TurboQuant-MSE)可降低所有测试协议中的悬崖,将无检测损失的工作点提升至7.9倍。该方法的步骤为:测量目标部署场景的悬崖,然后在其上方进行插值。

英文摘要

The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize. SemKV preserves every token, ranks tokens by a model-internal score, and assigns two adjacent above-cliff precisions, achieving a measured 6.0x storage reduction with no statistically detectable quality difference from full KV (n=900, three seeds), and outperforming FP16 token pruning granted a 1.5x larger memory budget. Replacing the affine base with a distortion-optimized quantizer (TurboQuant-MSE) lowers the cliff in every protocol tested, raising the no-detectable-loss operating point to 7.9x. The recipe: measure the cliff for the target deployment setting, then interpolate above it.

发表机构

  • Electronics and Telecommunications Research Institute (ETRI)(电子通信研究院(ETRI))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑