发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长思维链推理中KV缓存压缩,提出BreadthKV方法,通过量化与驱逐结合,在固定字节预算下优先增加缓存令牌数量而非精度,经60问题校准位宽,在多个基准上优于仅驱逐和现有方法。
AI 中文摘要
推理模型在解码长思维链(CoT)时,大部分KV缓存被写入,因此缓存必须在固定内存预算下在线压缩。解码时的方法主要决定要驱逐哪些令牌。我们研究固定字节预算应如何在缓存令牌数量及其精度之间分配。BreadthKV将字节花在更多低精度令牌上,结合量化和驱逐,并通过60个问题的端到端校准为每个模型和预算选择位宽,因为离线注意力误差无法可靠预测。在三个推理模型和四个数学与科学基准上,它在18个设置中的17个中得分高于仅驱逐,并产生更短的输出。驱逐损失的大部分来自脱轨运行,这些运行在未达到答案的情况下持续推理直到长度上限。在Qwen3-8B上,在我们最紧的预算下,驱逐将91%的AIME样本送到上限,而BreadthKV为40%。在相同协议下,BreadthKV与使用27%更多KV内存时间的联合率失真分配器(RDKV)在统计上无差异,并且优于我们重新实现的ThinKV。
英文摘要
Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses comes from derailed runs, which keep reasoning until the length cap without reaching an answer. On Qwen3-8B at our tightest budget, eviction sends 91% of AIME samples to the cap and BreadthKV 40%. Under the same protocol, BreadthKV is statistically indistinguishable from a joint rate-distortion allocator (RDKV) that uses 27% more KV memory-time, and it outperforms our re-implementation of ThinKV.
Comments17 pages, 4 figures, 13 tables