KVShareArena:跨上下文和模型检查点的KV缓存复用
KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
KVShareArena提出跨上下文和模型检查点的KV缓存复用基准,评估位置修正等方法,发现多来源场景下需重编码或训练,并发布工具包与排行榜。
AI中文摘要:
LLM服务系统已经能够复用KV缓存,但仅限于被复用的文本位于提示词最开头的情况。两种日益增长的工作负载打破了这一条件:检索增强生成服务器为每个查询组装不同的检索块集合,而多智能体协调器则阅读其他智能体撰写的报告。在新提示词中复用缓存时,缓存携带错误的位置信息,且从未关注过其他来源。缓存也可能由同一模型家族的不同检查点写入,这改变了存储的值。针对此类缓存的修复方法已在三个独立社区中出现,各自按自身标准衡量,而现有基准仅测试精确前缀复用(此场景下无信息损失)。KVShareArena在检索块和智能体报告上,跨提示词上下文和模型检查点对KV缓存复用进行基准测试。它通过每个方法在无缓存与完全重计算之间恢复差距的比例来评分,并计入使用缓存时的计算、内存和每请求延迟,同时单独报告构建缓存的一次性成本。我们发现,修正位置(无需重计算)在问题需要同时涉及多个来源之前已足够。而在那之后,只有通过部分重编码缓存或训练来付出代价的方法,才能恢复差距的一半到三分之二;未修复的缓存可能比无缓存更差。在单个提示词上无害的缓存压缩方法,在新鲜写入的智能体报告上明显落后于位置修正。这些模式在三个模型板上保持一致。当不同检查点写入缓存时,免训练方法几乎不受影响,而针对一个检查点缓存训练的自适应器则质量下降。工具套件、冻结查询集和成本核算以pip包形式发布,附带自动化提交流程和公开排行榜。
英文摘要:
Reusing key-value (KV) caches speeds up LLM inference by avoiding repeated computation on shared text. Standard prefix caching reuses a KV cache only when the LLM is the same and all preceding text is identical, but real workloads often break both conditions: RAG systems place different documents before the same one, agents with different system prompts read the same file or tool output, multi-agent workflows use specialized LLMs on shared material, and an updated model reads documents cached by its previous version. Because KV caches depend on both the preceding text and the model weights, direct reuse can reduce answer quality. Many methods repair or compress the reused cache, but each paper uses its own tasks, models, and cost measures, and existing benchmarks mainly test long-context processing or reuse of an unchanged prefix. We introduce KVShareArena, a benchmark and open evaluation framework for comparing them under the same conditions. KVShareArena has (1) reuse tests on 2,150 questions from three QA datasets, where the preceding text, the cache-writing LLM, or both change while the answering LLM and input stay fixed; (2) five dense and mixture-of-experts LLMs (4B-30B) and six LLM pairs where one version of an LLM reads caches written by another, for 33 model-dataset settings; (3) 11 repair and compression methods from six method classes; (4) four evaluation perspectives: answer quality, prefill computation, KV-cache memory, and latency; and (5) a common interface for adding new methods and an interactive leaderboard. Experiments yield two findings. First, both the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models. Second, most repairs keep their quality when another LLM version wrote the cache, but a trained repair adapter loses quality in 12 of 18 pair-dataset tests. Code and data: https://github.com/xishi404/KVShare-Arena