arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BeaconKV:由信标查询引导的键值缓存压缩,用于实现高效的大型推理模型推理

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi

arXiv 2609.04971首次发表:更新:

发表机构

Hanyang University; Sungkyunkwan University(汉阳大学; 成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BeaconKV是一种无需训练的KV缓存压缩方法,通过维护代表查询聚类的信标查询预测需重访的KV对,在4种开源LRMs上实现最高5.8倍内存缩减,同时保留高准确率并提升4.3倍以上吞吐量,性能优于现有方法。

AI 中文摘要

大型推理模型(Large Reasoning Models, LRMs)通过扩展思维链(Chain-of-Thought, CoT)生成实现了卓越的问题解决能力,但由此产生的键值(Key-Value, KV)缓存会随序列长度线性增长,造成严重的内存瓶颈,对于长推理轨迹而言往往会超出GPU的容量。现有的KV缓存压缩方法依赖近期查询来估计未来token的重要性,默认假设这些查询可作为未来注意力模式的可靠代理。我们证明该假设在长程推理中不成立:某些解码步骤会生成思维重访token(Thought Revisiting Tokens, TRT),这些token会重新关注之前的远距离上下文,例如推理轨迹早期制定的任务解决计划。通过系统分析,我们发现与TRT对应的查询在嵌入空间中会聚类为少量相似组。基于这一见解,我们提出了BeaconKV,这是一种无需训练的KV缓存压缩方法,它维护信标查询(即每个全局查询聚类的紧凑代表),以预测哪些KV对将被重访,而无需存储整个查询历史。在四个开源LRMs和多种推理基准上,BeaconKV的表现总体优于现有压缩方法,实现了高达5.8倍的内存减少,同时几乎保留了完整的缓存准确率,并将吞吐量提高了4.3倍以上。

英文摘要

Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to $5.8\times$ memory reduction while nearly preserving full cache accuracy and improving throughput by over $4.3\times$.

CommentsICML 2026. Code: https://github.com/aiha-lab/BeaconKV

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑