AI 中文总结
提出草稿引导驱逐(DGE),通过延迟驱逐至草拟前两个答案令牌后,解决免训练KV缓存压缩中驱逐时机问题,在LongBench上达到44.2,接近FullKV。
AI 中文摘要
免训练的KV缓存压缩方法(如SnapKV、H2O和PyramidKV)在预填充阶段结束时驱逐令牌,旨在保留未来查询预期使用的注意力质量——优化保留什么。我们表明这一目标在两个方面失败。(1)补偿:恢复被驱逐的注意力质量可以恢复注意力级别的目标,而无需恢复任务质量。(2)选择:覆盖更多真实的解码查询质量可能损害质量,当恢复的质量是碎片化的而非集中在连贯的片段中时。这些失败有一个共同原因:驱逐发生在决定答案轨迹的查询存在之前。我们提出草稿引导驱逐(DGE),它将驱逐推迟到使用完整缓存起草前k=2个答案令牌之后——仅比预填充多一个解码步骤。由于草稿是从答案自身的前缀生成的,没有缓存条目在此轨迹信号可用之前被丢弃。每头缓存预算保持不变,DGE可以直接应用于SnapKV、PyramidKV、H2O和StreamingLLM,而无需修改它们的驱逐分数。与额外通行方法不同,DGE改变驱逐发生的时间,而不是选择哪些缓存条目。大量实验表明,在六个指令调优骨干中的五个上,DGE在每个评估预算下都优于先前方法,在LongBench上达到44.2,几乎匹配FullKV的44.3。仅时序控制的DGE-W达到相同分数,表明增益来自驱逐发生的时间而非选择的内容——我们将此效应称为轨迹锚定。
英文摘要
Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attention mass can recover the attention-level target without recovering task quality. (2) Selection: covering more of the true decode-query mass can hurt quality when the recovered mass is fragmented rather than concentrated in coherent spans. These failures share a common cause: eviction occurs before the queries that determine the answer trajectory exist. We propose Draft-Guided Eviction (DGE), which defers eviction until after drafting the first k=2 answer tokens using the full cache - just one decode step beyond prefill. Because the draft is generated from the answer's own prefix, no cache entries are discarded before this trajectory signal becomes available. The per-head cache budget remains unchanged, and DGE can be applied directly to SnapKV, PyramidKV, H2O, and StreamingLLM without modifying their eviction scores. Unlike extra-pass methods, DGE changes when eviction occurs rather than what cache entries are selected. Extensive experiments demonstrate that DGE outperforms prior methods at every evaluated budget on five of six instruct-tuned backbones, achieving 44.2 on LongBench, nearly matching FullKV at 44.3. The timing-only control DGE-W achieves the same score, demonstrating that the gain comes from when eviction occurs rather than what is selected - an effect we term trajectory anchoring.