AI 中文总结
QEvict是一种三层KV缓存管理方案,通过可恢复量化驱逐解决长上下文解码中的注意力漂移问题,在固定内存预算下提升了长上下文任务的信息保留效果。
AI 中文摘要
自回归大语言模型推理越来越受到键值(KV)缓存内存占用的限制。主流工作通过驱逐基于注意力衍生分数被认为不重要的token来降低该占用。然而,这类策略做出了一个隐含的不可逆决策:token一旦被驱逐,就无法再次发挥作用。我们证明该假设在解码过程中是脆弱的。随着生成的查询不断演变,token和窗口的重要性会发生漂移,导致标准驱逐策略永久丢弃那些在全缓存模型下后续会受到大量注意力关注的状态。为表征该行为,我们引入了Future Missed Mass和Global LIR这两个诊断指标,分别用于测量分配给被丢弃状态的未来注意力,以及历史非活跃区域的重新激活情况。我们提出QEvict,这是一种三层KV缓存管理方案,它将二元保留或删除的驱逐替换为可恢复驱逐。QEvict以全精度维护高置信度窗口,将中间窗口存储在量化可恢复层中,仅删除置信度最低的窗口。解码过程中,累积注意力分数会更新窗口重要性,当量化窗口再次变得重要时,它会被反量化并提升至全精度层。在固定内存预算下,该设计在保留最重要区域精确全精度的同时,保留了更广泛的历史上下文。在长上下文理解、检索和推理基准测试中,QEvict始终优于代表性的驱逐和量化基线,减少了注意力遗漏并提升了信息保留效果。
英文摘要
Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evicting tokens that appear unimportant under attention-derived scores. However, such policies make an implicit irreversible decision: once a token is evicted, it cannot become useful again. We show that this assumption is brittle during decoding. Token and window importance drift as generated queries evolve, causing standard eviction policies to permanently discard states that later receive substantial attention under the full-cache model. To characterize this behaviour, we introduce Future Missed Mass and Global LIR, two diagnostics that measure future attention assigned to discarded states and the reactivation of historically inactive regions. We propose QEvict, a three-tier KV-cache management scheme that replaces binary retain-or-delete eviction with recoverable eviction. QEvict maintains high-confidence windows in full precision, stores intermediate windows in a quantized recoverable tier, and deletes only the lowest-confidence windows. During decoding, cumulative attention scores update window importance and when a quantized window becomes important again, it is dequantized and promoted to the full-precision. Under a fixed memory budget, this design preserves broader historical context while retaining exact full precision for the most important regions. Across long-context understanding, retrieval, and reasoning benchmarks, QEvict consistently improves over representative eviction and quantization baselines, reducing missed attention and improving information retention
Comments24 pages, 6 figures. The first four authors contributed equally