arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15498cs.CLcs.LG

VarRate:用于长上下文语言模型的免训练可变率键值缓存压缩

VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs

Shahrzad Esmat, Dhawal Shah, Ali Jannesari

首次发表
浏览论文内容

中文总结 AI 辅助

研究长上下文LLM推理中KV缓存内存瓶颈问题,提出免训练的VarRate编解码器,依查询显著性为令牌分配可变低秩预算,无需训练,相比其他方法在准确率和开销上表现出色,是较强的匹配内存压缩器。

中文摘要 AI 辅助

键值(KV)缓存是长上下文大语言模型(LLM)推理中的主要内存瓶颈。两种主要的免训练方法在结构上都有局限:令牌选择方法(SnapKV、Ada-KV)从观察窗口评估重要性并逐出低分令牌,但逐出不可逆,在查询无关重用下重要性信号退化时,准确率会下降11 - 15分;均匀低秩编码保留每个令牌但各处秩相同,浪费预算。我们发现两种方法的问题都可用分配秩而非逐出解决。我们提出VarRate,一种免训练的KV编解码器,根据查询显著性为每个令牌分配可变低秩预算,使每个令牌秩非零。可比的自适应秩编解码器需训练,VarRate无需。因无令牌被丢弃,在查询感知选择崩溃的情况下,它仅下降3.5 - 5.5分。在LongBench(16个任务)上20%预算匹配时,VarRate在Llama-3.1-8B和Qwen2.5-7B上与未压缩模型相差不超0.8分。平均而言,它是最强的匹配内存压缩器,显著优于均匀秩消融模型,在与专为查询无关重用设计的KVzip对比中,在四种设置中的三种下准确率相当,总体相差不超一分,预填充开销约为其八分之一。

英文摘要

The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible -- so when the importance signal degrades under query-agnostic reuse, accuracy collapses by 11-15 points; uniform low-rank coding keeps every token but spends equal rank everywhere, wasting budget. We observe that both failures share one cure: rank should be allocated, not evicted. We present VarRate, a training-free KV codec that assigns each token a variable low-rank budget by its query salience, keeping every token at a nonzero rank. Comparable adaptive-rank codecs reach this allocation only through training; VarRate requires none. Because no token is dropped, it degrades by only 3.5-5.5 points where query-aware selection collapses. At a matched 20% budget on LongBench (16 tasks), VarRate stays within 0.8 points of the uncompressed model on both Llama-3.1-8B and Qwen2.5-7B. Averaged over the two, it is the strongest matched-memory compressor. It significantly beats its uniform-rank ablation on both models. Against KVzip, a method purpose-built for query-agnostic reuse, it is accuracy-equivalent in three of four settings and within a point overall, at about one-eighth the prefill overhead.

补充信息

↑