arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DistillCache:基于KL引导的自适应KV缓存驱逐的内存高效大语言模型推理

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

Asaad Althoubi

arXiv 2608.08878首次发表:更新:

发表机构

Oklahoma State University(俄克拉荷马州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DistillCache是一种强化学习框架,将KV缓存驱逐建模为序列决策问题,在25%缓存预算下,在LongBench上保留全缓存94.2%的准确率,优于多种基线方法,可提升推理吞吐量。

AI 中文摘要

基于Transformer的大语言模型(LLM)在众多任务中表现出色,但其键值(KV)缓存会随序列长度线性增长,给长上下文推理带来严重的内存瓶颈。现有的启发式驱逐方法(如H₂O和SnapKV)依赖静态注意力或位置信号,往往无法捕获token的未来预测影响力。我们提出DistillCache,这是一个将KV缓存驱逐建模为序列决策问题的强化学习框架。DistillCache利用丰富的内部模型信号(注意力统计量、值范数、熵和位置)学习轻量级策略网络,并通过每步KL散度奖励使用REINFORCE算法进行训练,以保留全缓存的输出分布。在7B参数的指令调优Transformer(Mistral-7B-Instruct-v0.3)上,当缓存预算为25%时,DistillCache在LongBench上保留了全缓存94.2%的准确率,比强大的启发式基线(H₂O、SnapKV)高出最多2.7个绝对点;在我们的重新实现中,比并行的基于强化学习的方法(ForesightKV、RLKV)在长上下文任务上高出最多1.4个点。在推理基准上,DistillCache与最佳并行方法具有竞争力,且在激进压缩下超过该方法。它还提供了最多2.1倍的全缓存吞吐量,同时保持有竞争力的实际效率。这些结果凸显了学习到的、感知分布的策略对于内存高效长上下文LLM推理的有效性。

英文摘要

Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for long-context inference. Existing heuristic eviction methods (e.g., H$_2$O and SnapKV) rely on static attention or positional signals that often fail to capture a token's future predictive influence. We propose DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem. DistillCache learns a lightweight policy network using rich internal model signals (attention statistics, value norms, entropy, and position) and trains it with REINFORCE via a per-step KL-divergence reward to preserve the full-cache output distribution. On a 7B-parameter instruction-tuned Transformer (Mistral-7B-Instruct-v0.3), DistillCache retains 94.2% of full-cache accuracy on LongBench at a 25% cache budget, outperforming both strong heuristic baselines (H$_2$O, SnapKV) by up to 2.7 absolute points and, under our re-implementations, concurrent RL-based methods (ForesightKV, RLKV) by up to 1.4 points on long-context tasks. On reasoning benchmarks, DistillCache is competitive with the best concurrent method and surpasses it under aggressive compression. It also delivers up to 2.1x full-cache throughput while maintaining competitive practical efficiency. These results highlight the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Comments20 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑