arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARC-KV:基于重建的KV缓存压缩中的锚点搜索摊销

ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li

arXiv 2609.36835首次发表:更新:

发表机构

University of Maryland, College Park; Meta(马里兰大学学院公园分校; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ARC-KV提出选择性摊销原则,训练值感知索引器单次选择锚点并复用,结合凸包键合并与注意力质量偏置,在长上下文KV压缩中显著提升效率与准确率。

AI 中文摘要

长上下文大语言模型推理受到KV缓存的瓶颈限制,该缓存随序列长度线性增长。这一负担对于长且可复用的上下文前缀尤为严重,因为其缓存必须服务于许多下游查询。基于重建的方法(如Attention Matching)通过紧凑的KV缓存实现了强大的下游任务性能。然而,迭代锚点搜索主导了基于OMP的Attention Matching的压缩成本。这促使我们提出选择性摊销原则,即跨上下文学习可复用的锚点选择策略,同时保留特定于上下文的重建。在这项工作中,我们提出了ARC-KV,一种遵循该原则的新型基于重建的KV缓存压缩方法。为此,我们首先训练一个值感知索引器,在单次评分过程中选择真实键锚点。然后,ARC-KV应用凸包约束的键合并,并针对完整缓存拟合注意力质量偏置和紧凑值。在推理时,ARC-KV使用冻结的索引器为每个上下文构建一次紧凑缓存,并将其复用于所有后续查询。大量实验表明,在Llama-3.1-8B-Instruct上的QuALITY、RULER和LongBench中,ARC-KV在大多数设置下优于已报道的压缩方法。特别是,在QuALITY上保留10% KV时,ARC-KV将准确率从0.6409提高到0.6474,相较于Attention Matching,同时将压缩时间从959.8秒减少到37.3秒,降低了25.73倍。

英文摘要

Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑