发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 KV 缓存重练的主流方法在严格预算下效果不佳,本文提出 Attic-KV 方法,通过引用上下文生成问答对自我测试并结合锚定标记,在多个基准任务中显著提升了长上下文处理性能。
AI 中文摘要
许多键值(KV)缓存会在明确其用途前就被压缩:用于检索的文档、跨请求共享的提示前缀、长对话的记忆等。主流方法通过重练对 KV 条目评分:模型重读上下文并保留其关注的条目,假设缓存重练上下文越完整,记忆效果越好。我们证明在严格预算下该假设会适得其反:重练所有,不记任何。在 3% 保留率下,重读整个上下文在 RULER 上保留 96.5 分中的 31.5 分,在 LongBench 的自然文本任务中,其表现甚至低于完全不重练的方法。原因在于缓存保留其重练的内容:重读会将预算分散到整个上下文,导致答案自身的条目仅以略高于随机的概率留存。如同考试前的学生,缓存通过自我测试而非重读能记住更多。由此得出两个原则:重练将被读取的内容,且重练尽可能多的内容。我们将其实例化为 Attic-KV(简称 Attic),这是一种无需训练的重练方法,其中模型会引用上下文生成问答对进行自我测试,并结合内容自适应数量的锚定标记。仅改变重练方式,Attic 在 RULER 和 LongBench 的自然文本任务的 8 个测试设置中,是表现最佳的无需训练方法;将其接入基于梯度的 KVgrad 和已训练的 RestoreKV+,可分别将其性能提升最高 17.1 和 28.1 分。其优势随预算缩减而增大,在 3% 保留率下较全重读高出 41.9 分,且压缩速度快于重读整个上下文。
英文摘要
Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it. We show that under tight budgets this assumption backfires: rehearse everything, remember nothing. At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all. The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance. Like a student before an exam, a cache remembers more by testing itself than by rereading. Two principles follow: rehearse what will be read, and rehearse as much as there is. We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount. Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points. Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.
Comments14 pages, 5 figures