发表机构
UC San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WakeKV提出响应式可逆KV驻留策略,将冷却注意力头移至CPU存储池,在匹配预算下优于冻结分类和破坏性驱逐,提升吞吐量并保持质量。
AI 中文摘要
大多数KV缓存压缩方法只对注意力头进行一次分类,要么在离线阶段,要么在预填充阶段,并在整个生成过程中保持这一分类不变。在三个模型(1.5B-8B)和三种场景(针检索、长思维链和多轮召回)中,我们在四种模型-场景组合上测量了注意力头的行为,发现大多数注意力头在生成过程中至少改变一次其读取行为。我们提出WakeKV,一种响应式驻留策略,将冷却的注意力头移动到可恢复的CPU存储池,而不是冻结或永久驱逐其状态。在匹配的内存或预算下,WakeKV在五种模型-场景组合上始终优于冻结分类和破坏性驱逐的未命中率,并在四种符合条件的组合上优于三个引用的基线(SnapKV、uniform R-KV和ReasonAlloc)。在Mistral-7B上的FlexiCache/vLLM实现证实了在真实硬件上的优势,提高了吞吐量,同时保持了LongBench质量。
英文摘要
Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.
CommentsAccepted to the NeurIPS 2026 Workshop on ML for Systems. 2 figures, 4 tables, appendix