发表机构
Duksung Women’s University; KAIST(德成女子大学; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨上下文KV缓存复用导致准确率下降的问题,提出预算化缓存修复(BCR),通过草稿token注意力排序并重算固定预算的缓存行,在保持缓存服务的同时恢复密集预填充准确率。
AI 中文摘要
跨上下文KV缓存复用预测共享段在新前缀下的键和值,而非重新计算它们,且已有报告称这样做不会损失质量。我们发现并非如此,并识别出两个问题。(1)隐藏成本:在MMLU和GSM8K上,复用会损失大量准确率。(2)决策单位错误:没有决定是否复用缓存的规则能消除该成本。真正有帮助的是选择缓存中哪些部分需要重新计算,而选择得当的价值随选择单位增大而下降:在单行(一个token的键和值)上进行明智选择可消除超出随机水平的49.5%的缓存误差,在64-token块上为10.6%,而在整个调用层面则毫无效果。预算化缓存修复(BCR)在选择仍有价值的单位上运作。它从组装好的缓存中草拟两个token,根据这些token对缓存行的注意力对行进行排序,并精确地以三种布局之一重新计算固定预算的行。成本是实际支付而非预测规避,而作为门控失败的草稿却作为选择器成功。BCR将GSM8K恢复到密集预填充的准确率,同时仍从缓存服务大多数调用,其最佳布局在参考网格中优于所有复用基线的均值。草稿还在相同预算下击败了抛硬币选择器——这是先前评估所缺乏的对照。
英文摘要
Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.