arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07643cs.CLcs.AI

KV缓存驱逐的蒙特卡洛估计

Monte Carlo Estimation for KV Cache Eviction

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Wajih Hassan Raza, Atta Ul Asad, Young D. Kwon, Michal Valko, Dean F. Hougen

首次发表
浏览论文内容

中文总结 AI 辅助

提出LORE-KV,一种免训练的蒙特卡洛方法,通过采样未来查询轨迹估计令牌效用,实现KV缓存驱逐,在LongBench和RULER上显著提升性能。

中文摘要 AI 辅助

大多数KV缓存驱逐方法实际上是在问:在阅读提示时,哪些记忆显得重要?而我们问的是:在回答时,哪些记忆将起作用?由于解码查询在驱逐时不可用,先前具有未来感知的方法依赖于伪响应或合成未来查询估计。我们将固定预算的未来感知驱逐视为对合理的模型条件查询轨迹的分布估计,并引入LORE-KV(基于可靠性加权集成的前瞻输出扰动,用于键值缓存),这是一种免训练方法,它从冻结的目标模型中采样短的自回归延续,并使用其响应侧查询状态来估计提示令牌的效用。令牌通过投影的留一法注意力输出删除成本进行评分,并在采样未来中通过可选的轨迹加权进行聚合。临时延续在最终解码前被丢弃,无需辅助模型或训练。消融实验隔离了该机制:在B=128时,单个响应侧延续恢复了相对于提示窗口控制约89%的增益,而额外的未来提供较小的改进。在B=128时,LORE-KV将Qwen2.5-14B上的LongBench平均值从45.49提高到48.24(+2.75),并将Mistral-7B上的16K RULER平均值从45.20提高到51.05(+5.85)。在较大的缓存预算下,增益减少,并与任务级回归共存。LORE-KV在六个密集和混合注意力骨干上,作为一次性压缩开销,其每样本墙钟时间为AnDPro的1.46-2.77倍。

英文摘要

Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected leave-one-out attention-output deletion cost and aggregated across sampled futures with optional trajectory weighting. The temporary continuations are discarded before final decoding, requiring no auxiliary model or training. Ablations isolate the mechanism: at B=128, a single response-side continuation recovers about 89% of the gain over the prompt-window control, while additional futures provide smaller improvements. At B=128, LORE-KV raises the LongBench average on Qwen2.5-14B from 45.49 to 48.24 (+2.75) and the 16K RULER average on Mistral-7B from 45.20 to 51.05 (+5.85). Gains diminish at larger cache budgets and coexist with task-level regressions. LORE-KV incurs 1.46-2.77x AnDPro's per-sample wall-clock time as a one-time compression overhead across six dense and hybrid-attention backbones.

发表机构

  • University of Oklahoma(俄克拉荷马大学)
  • Stanford University(斯坦福大学)
  • University of Houston(休斯顿大学)
  • Lahore University of Management Sciences(拉合尔管理科学大学)
  • University of Cambridge(剑桥大学)
  • Isara Labs(Isara 实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑