通过大语言模型引导的程序演化发现KV缓存驱逐策略
Discovering KV Cache Eviction Policies via LLM-Guided Program Evolution
AI总结:
本文提出CacheCraft方法,通过大语言模型引导的程序演化自动发现KV缓存驱逐策略,得到FRC评分器,在多模型多压缩率下性能优于基线,还提供了可迁移的自动策略发现方案。
AI中文摘要:
KV缓存压缩对长上下文推理至关重要,但设计有效的驱逐策略仍存在困难:现有预填充阶段方法通常依赖手动设计的显著性启发式规则,这些规则在不同模型、上下文长度和压缩率下的鲁棒性较差。本文提出CacheCraft,一种使用大语言模型引导的代码演化引擎自动发现KV缓存驱逐策略的程序演化方法。CacheCraft发现了FRC(特征丰富压缩),这是一种固定权重的三信号评分器,结合了接收的局部注意力、邻域注意力密度以及KV头最大显著性,并采用分块级top-k选择。无需针对每个模型进行重新调整,FRC在所有RULER 4k/8k单元(r≥0.75)的评估单通KVPress基线中排名第一,涉及Llama-3.1-8B-Instruct和Qwen3-8B的20个网格单元中的12个,在88%压缩率下,Llama-4k的性能提升15.4个点,Qwen-8k提升13.9个点。评分器与结构的分解实验显示,评分器家族而非分块选择是关键设计选择:引入评分器贡献了67.2个RULER点,而改进分块结构仅贡献约0.1个点。除FRC本身外,CacheCraft还提供了一种可迁移的自动驱逐策略发现方案:紧凑的策略接口、具有严格输出不变量的级联评估器,以及将搜索平台和奖励作弊失败视为重构可编辑接口证据的诊断循环。
英文摘要:
KV cache compression is critical for long-context inference, yet effective eviction policies remain difficult to design: existing prefill-stage methods often rely on hand-crafted salience heuristics that can be brittle across models, context lengths, and compression ratios. We present CacheCraft, a program-evolution methodology for automatically discovering KV cache eviction policies using an LLM-guided code-evolution engine. CacheCraft discovers FRC (Feature-Rich Compression), a fixed-weight three-signal scorer that combines local attention received, neighborhood attention density, and KV-head maximum salience with chunk-level top-k selection. Without per-model retuning, FRC ranks first among the evaluated single-pass KVPress baselines at every RULER 4k/8k cell with r >= 0.75 across Llama-3.1-8B-Instruct and Qwen3-8B (12 of 20 grid cells), gaining +15.4 points on Llama-4k and +13.9 points on Qwen-8k at 88% compression. A scorer-versus-structure decomposition shows that the scoring family, not chunk selection, is the load-bearing design choice: incorporating the scorer contributes +67.2 RULER points, while improving chunk structure contributes only ~0.1. Beyond FRC itself, CacheCraft provides a transferable recipe for automated eviction-policy discovery: a compact policy interface, a cascade evaluator with strict output invariants, and a diagnostic loop that treats search plateaus and reward-hacking failures as evidence for reformulating the editable interface.