AvoKV-E:面向长推理的负载感知KV缓存驱逐
AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning
- LinkedIn Corporation(领英公司)
- Clemson University(克莱姆森大学)
- Rice University(莱斯大学)
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AvoKV-E提出无需训练的负载感知KV缓存驱逐策略,通过延迟资格、归一化读取压力及值负载潜力排序,在严格缓存下超越现有基线,强调保留值负载以维持长推理轨迹。
AI中文摘要:
长输出推理将KV缓存瓶颈从固定提示转移到生成的轨迹上。现有的推理缓存驱逐方法大多将缓存条目视为路由对象,估计旧键是否仍会被读取、是否会重现或可被替换。这种仅基于路由的视角忽略了两方面的影响:低注意力条目可能携带较大的值负载,其移除会改变未来预测;新生成的状态可能在后续查询有机会读取之前就显得陈旧。我们提出AvoKV-E,一种无需训练的驱逐策略,首先延迟近期状态的资格,然后使用候选归一化读取压力、键冗余和值负载潜力对合格条目进行排序。根据跨不同模型和数据集的实证评估,在匹配的活动KV预算下,AvoKV-E达到或超过冗余感知、基于重现和思维自适应的驱逐基线,其最大增益出现在最严格的缓存场景中。组件和反事实分析进一步将这些增益与延迟观察、负载感知评分、冗余和尺度鲁棒归一化联系起来。总体结果表明,长推理KV驱逐不仅应保留可能被读取的键,还应保留维持推理轨迹的值负载。
英文摘要:
Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.