arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03007cs.LGcs.AI

AvoKV-E:面向长推理的负载感知KV缓存驱逐

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

  • LinkedIn Corporation(领英公司)
  • Clemson University(克莱姆森大学)
  • Rice University(莱斯大学)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hao Wang, Alborz Geramifard

AI总结:

AvoKV-E提出无需训练的负载感知KV缓存驱逐策略,通过延迟资格、归一化读取压力及值负载潜力排序,在严格缓存下超越现有基线,强调保留值负载以维持长推理轨迹。

AI中文摘要:

长输出推理将KV缓存瓶颈从固定提示转移到生成的轨迹上。现有的推理缓存驱逐方法大多将缓存条目视为路由对象,估计旧键是否仍会被读取、是否会重现或可被替换。这种仅基于路由的视角忽略了两方面的影响:低注意力条目可能携带较大的值负载,其移除会改变未来预测;新生成的状态可能在后续查询有机会读取之前就显得陈旧。我们提出AvoKV-E,一种无需训练的驱逐策略,首先延迟近期状态的资格,然后使用候选归一化读取压力、键冗余和值负载潜力对合格条目进行排序。根据跨不同模型和数据集的实证评估,在匹配的活动KV预算下,AvoKV-E达到或超过冗余感知、基于重现和思维自适应的驱逐基线,其最大增益出现在最严格的缓存场景中。组件和反事实分析进一步将这些增益与延迟观察、负载感知评分、冗余和尺度鲁棒归一化联系起来。总体结果表明,长推理KV驱逐不仅应保留可能被读取的键,还应保留维持推理轨迹的值负载。

英文摘要:

Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.

↑