发表机构
University of Notre Dame(圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长上下文LLM解码中KV缓存的内存瓶颈,提出与NVM接口协同设计的量化方案,通过固定码本和低元数据开销,在保持准确性的同时显著降低读取能量。
AI 中文摘要
键值(KV)缓存是长上下文大语言模型(LLM)解码中的主要内存瓶颈:每一步都需要完整读取它,因此解码受限于内存带宽。将量化的KV缓存保存在密集的片上非易失性存储器(NVM)中可消除片外传输。然而,现有的KV量化方法是为GPU式内存系统设计的:KIVI附加每组元数据,使存储的KV缓存增加约25%;KVQuant保留稀疏的全精度离群值,而密集阵列无法原位容纳这些值。本文研究了当KV缓存驻留在具有固定范围转换器的NVM中时这些结构的代价,并设计了一种与该接口匹配的量化方案。该架构将量化的KV缓存存储在密集的片上NVM中,仅使用一个小型静态模拟交叉阵列用于固定旋转,并将注意力计算保留在片上数字逻辑中。随机旋转和每向量归一化使每个坐标具有相同的范围,因此一个固定的键码本和一个固定的值码本(每个码本共享于对应张量类型的所有令牌)即可服务整个KV缓存。码本阈值被一次性编程为读取转换器的参考电平,实现固定范围数字化,无需逐令牌重新配置转换器。反量化是十六项查找和一次范数乘法;唯一的每向量元数据是一个标量,约占3%。在3B到14B的模型和长达32k令牌的上下文中,四比特KV缓存在存储和交叉阵列噪声(在现实器件水平下模拟)下保持准确性。KIVI和KVQuant在软件中仍更准确;我们格式的优势在于内存接口:映射到相同NVM时,KV读取能量比两者低3.1-3.6倍,元数据开销比KIVI低8倍。贡献在于与NVM内存接口协同设计的KV量化,而非新的准确率记录。
英文摘要
The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound. Holding a quantized KV cache in dense on-chip non-volatile memory (NVM) removes the off-chip transfer. Existing KV quantization methods, however, were designed for GPU-style memory systems: KIVI attaches per-group metadata, adding about 25% to the stored KV cache; KVQuant keeps sparse full-precision outliers that a dense array cannot hold in place. This paper examines what these structures cost when the KV cache resides in NVM behind fixed-range converters, and designs a quantization scheme matched to that interface. The architecture stores the quantized KV cache in dense on-chip NVM, uses a small static analog crossbar only for the fixed rotation, and keeps attention in on-chip digital logic. A randomized rotation and per-vector normalization give every coordinate the same range, so one fixed codebook for keys and one for values, each shared across all tokens of the corresponding tensor type, serve the entire KV cache. The codebook thresholds are programmed once as the read converter's reference levels, enabling fixed-range digitization with no per-token converter reconfiguration. Dequantization is a sixteen-entry lookup and one norm multiply; the only per-vector metadata is one scalar, about 3%. Across models from 3B to 14B and contexts to 32k tokens, the four-bit KV cache maintains accuracy under storage and crossbar noise simulated at realistic device levels. KIVI and KVQuant remain more accurate in software; the advantage of our format lies at the memory interface: 3.1-3.6x lower KV read energy than both mapped to the same NVM, and 8x lower metadata overhead than KIVI. The contribution is a KV quantization co-designed with the NVM memory interface rather than a new accuracy record.
CommentsIEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)