arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

激进解码时KV驱逐的关键因素:时间聚合与排名保留

What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

Bo Zeng, Yu Zhao, Yefeng Liu, Zhihong Lu, Xuanfan Ni, Xintong Wang

arXiv 2609.03515首次发表:更新:

发表机构

Ant Group; Alibaba Group; Universität Hamburg(蚂蚁集团; 阿里巴巴集团; 汉堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究聚焦激进解码时KV驱逐的关键因素,提出基于EMA的InertiaKV及其变体,揭示时间聚合与排名保留的重要性,相关方法可提升解码吞吐量且不显著降低质量。

AI 中文摘要

解码时KV缓存压缩研究主要聚焦于设计更优的令牌评分函数,而跨解码步骤聚合评分的时间规则常被视为实现细节。在激进KV压缩下,我们发现指数移动平均(EMA)聚合会使近似保序的评分器修改在驱逐集层面几乎无法区分;值范数与熵变体仍与注意力高度相关,产生几乎不变的保留集,而KeyDiff、键范数、近因性及学习型评分器会改变排名并大幅退化。我们将这种稳定性与所评估的聚合关联,该聚合耦合了层权重与时间保留。基于此观察,我们提出InertiaKV(一种基于EMA的解码时驱逐方法)及其定期刷新变体InertiaKV-Lazy,其解码吞吐量是全刷新InertiaKV的1.34-1.46倍。我们还将无评分解码作为独立经验操作点研究:它在首个解码步骤对全上下文评分一次,冻结该排名,平均质量变化为+0.03,同时消除所有后续评分。在六个开放权重主干及LongBench、LongBench-v2、RULER基准上,结果表明时间聚合与排名保留是独特且重要的设计因素,并非意味着评分质量总体无关紧要。

英文摘要

Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponential-moving-average (EMA) aggregation makes approximately order-preserving scorer modifications largely indistinguishable at the eviction-set level. Value-norm and entropy variants remain highly correlated with attention and produce nearly unchanged retention sets, whereas KeyDiff, key norm, recency, and a learned scorer alter the ranking and degrade substantially. We associate this stability with the evaluated aggregation, which couples layer weighting and temporal retention. Building on this observation, we introduce InertiaKV, an EMA-based decoding-time eviction method, and InertiaKV-Lazy, its periodic-refresh variant, which yields 1.34-1.46x decode throughput relative to full refresh InertiaKV. We also study Score-Free decoding as a separate empirical operating point: it scores the full context once at the first decode step, freezes that ranking, and incurs an average quality change of +0.03 while removing all subsequent scoring. Across six open-weight backbones and the LongBench, LongBench-v2, and RULER benchmarks, the results identify temporal aggregation and ranking preservation as distinct, consequential design factors; they do not imply that scoring quality is irrelevant in general.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑