arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

读取关键信息:面向KV缓存的查询自适应量化

Read What Matters: Query-Adaptive Quantization for KV Caches

Siddharth Bhandari, Lucas Gretta, Krishna Balasubramanian, Shiva Kasiviswanathan

arXiv 2610.11245首次发表:更新:

发表机构

Amazon; University of California, Berkeley(亚马逊; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ReadKV方法,通过查询自适应分配KV缓存的渐进式编码前缀,在降低读取成本的同时保持模型性能,在长上下文任务上延迟显著低于TurboQuant。

AI 中文摘要

KV缓存条目在未来查询未知时就已存储,但每个解码查询需要不同位置的精度。我们通过为保留位和每次查询获取的位设置独立预算来研究这种不匹配。ReadKV将每个键和值存储在渐进式编码中,其前缀支持不同的重建精度。对于每个查询,它会利用该查询分配键通道前缀,从重建的键计算注意力,然后利用该注意力分配值标记前缀,存储的条目保持不变。每个阶段在固定预算下优化校准的失真目标;我们证明了在细化增益递减情况下的精确分配,并将这些目标与注意力输出误差关联起来。我们还展示了一个有限维注意力族,在相同读取预算下,依赖查询的访问严格优于所有不依赖查询的读取器,即使存在不受限制的竞争编码器和解码器。在六个基础模型上,从8位缓存中平均读取4位,C4困惑度最多增加0.66%,使用的逻辑读取约为16位缓存的四分之一,保留容量为其一半。在相同的有效载荷读取预算下,它始终比存储并完全读取4位更准确。保留比每次查询获取更多位的目的是针对长上下文解码,其中每步移动的缓存字节而非权重主导成本。在两个指令调优模型上进行的长上下文问答和检索提供了额外的质量证据。在NVIDIA A10G上测试的8K令牌、单批次、单层工作负载中,具有2位平均有效载荷读取预算的受限8位ReadKV读取器比测试的TurboQuant编解码器延迟低39%。

英文摘要

KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhibit a finite-dimensional attention family where query-dependent access strictly outperforms every query-independent reader at the same read budget, even with unrestricted competing encoders and decoders. Across six base models, reading four bits on average from an eight-bit cache increases C4 perplexity by at most 0.66%, using about one quarter of the logical reads and half the retained capacity of a 16-bit cache. It is consistently more accurate than storing and fully reading four bits at the same payload-read budget. Retaining more bits than each query fetches is aimed at long-context decoding, where the cache bytes moved per step, rather than the weights, dominate cost. Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence. On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.

Comments50 pages, 3 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑