发表机构
Hangzhou Institute for Advanced Study; University of Chinese Academy of Sciences(杭州高等研究院; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长上下文KV缓存压缩问题,提出CORE方法,通过覆盖校准和驱逐质量再分配实现高效缓存管理,在90%压缩率下显著提升性能。
AI 中文摘要
长上下文解码日益受到键值(KV)缓存内存和带宽的限制。现有的固定预算压缩方法通常将保留与补偿分开处理,而保留排序既未指定被丢弃的注意力质量,也未指定由此产生的输出误差方向。我们从精确分解出发:驱逐误差等于被驱逐的注意力质量乘以被驱逐质心与保留输出之间的方向差距,这凸显了保留中集合级覆盖以及保质量内存写入的重要性。我们提出了CORE(KV缓存的覆盖校准与驱逐质量再分配),该方法将结合查询效用和对数行列式覆盖的离线分配提炼为一个轻量级的缓存感知索引器。在推理时,一个校准的分布驱动两个通道:其Top-B排序保留互补的KV状态,而其排除的分配质量和条件权重则参数化潜在内存写入,无需单独的写入权重预测器或在线对数行列式评估。我们的分析提供了四项预补偿误差证书,刻画了非加性覆盖交互,并通过循环和查询自适应归一化建立了与质量无关的写入稳定性及分层界限。在三个骨干网络上,CORE在90%压缩率下比最强的RULER基线高出最多3.78个百分点;LongBench和重复驱逐评估进一步证明了其强大的有效性和解码效率。
英文摘要
Long-context decoding is increasingly constrained by key--value (KV) cache memory and bandwidth. Existing fixed-budget compression methods typically separate retention from compensation, while a retention ranking specifies neither discarded attention mass nor the direction of induced output error. We start from an exact factorization: eviction error equals evicted attention mass times the directional gap between the evicted centroid and retained output, highlighting the importance of set-level coverage in retention and mass-preserving memory writing. We introduce CORE COverage Calibration and Evicted-Mass REdistribution for KV Cache, which distills an offline allocation combining query utility and log-determinant coverage into a lightweight cache-aware indexer. At inference, one calibrated distribution drives both channels: its Top-$B$ ordering retains complementary KV states, while its excluded allocation mass and conditional weights parameterize latent-memory writes without a separate write-weight predictor or online log-determinant evaluation. Our analysis provides a four-term pre-compensation error certificate, characterizes non-additive coverage interactions, and establishes mass-independent write stability with hierarchical bounds through recurrent and query-adaptive normalization. Across three backbones, CORE exceeds the strongest RULER baseline by up to 3.78 points at 90\% compression; LongBench and repeated-eviction evaluations further demonstrate strong effectiveness and decoding efficiency.