发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究查询可见性对键值缓存压缩排名的影响,通过匹配预算审计六种压缩方法与三个基线,发现查询可见性改变排名,且方法在两种协议下的下降与问题对评分信号的可见度有关。
AI 中文摘要
键值缓存压缩方法主要在压缩前将查询附加到上下文中进行评估,这是一种查询感知协议。然而,压缩键值缓存的经济理由是重用:对文档进行一次压缩,然后针对它回答许多未来的问题。在这种部署中,压缩必须在查询不可知的情况下进行,即在看到任何问题之前。我们针对三个开放的7-9B模型上的三个简单基线,对六种已发表的压缩方法进行了匹配预算审计。所有因素保持不变,除了评分规则。有三个发现:(1)查询可见性会改变排名;(2)两种协议之间每种方法的下降与问题对每种方法评分信号的可见程度一致;
英文摘要
KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol. Yet the economic case for a compressed KV cache is reuse: compress a document once, answer many future questions against it. In that deployment, compression must happen query-agnostic -- before any question is seen. We present a matched-budget audit of six published compression methods against three trivial baselines on three open 7-9B models (144,300 paired evaluations on RULER-8192; 40,800 on LongBench; 50,000-resample paired bootstrap throughout). Everything is held fixed -- model, compression ratio, instances, decoding -- except the scoring rule. Three findings. (1) Query visibility changes the rankings: under the agnostic protocol, of the five audited methods that share a common attention backend, only KeyDiff beats a best-of-3 trivial baseline consistently (31 of 36 cells), and the most widely deployed method, SnapKV, loses to "keep the start and the recent window" on average (-0.066). (2) The per-method drop between the two protocols is ordered consistently with how visible the question is to each method's scoring signal, legible in its source code: from Delta=+0.198 for SnapKV (the question sits inside its 64-token observation window) down to Delta=+0.011 for KeyDiff (its score contains no query term at all).
Comments12 pages, 5 figures