发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究KV缓存逐出问题,确定性逐出存在误差估计不一致问题。提出随机逐出方法,通过泊松采样等技术实现可识别性和误差证书。在实际工作负载实验中,虽有声明未通过,但随机化带来了归因能力,能区分故障并更好安排重新计算。
AI 中文摘要
确定性KV缓存逐出会根据重要性分数保留前k个令牌并删除其余部分。我们证明这种设计无法知道它销毁了什么:被逐出的值可以被更改,使得服务系统保留的所有内容不变,而真实的注意力输出误差任意增长,因此该误差的任何服务时间估计器都不一致。随机逐出恢复了可识别性。通过在已知包含概率下进行泊松采样尾部,一个对数偏移在softmax内部执行Hájek校正,并且对保留集的调查采样方差估计器成为每步误差证书,经验覆盖率为0.97且无精度成本。在实际工作负载上,我们预先登记了七个声明,其中三个未通过:在25%-50%预算下的问题感知逐出几乎没有成本;输出对数概率比证书更好地预测失败;证书门控预算升级没有任何作用。幸存下来的是归因:证书将缓存引起的故障与固有故障分开(AUC为0.73-0.75,而输出置信度为0.47-0.54),并且比随机或置信门控更好地安排重新计算。随机化带来的是归因,而不是预测。
英文摘要
Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest, and after the deletion the serving system cannot know what the eviction cost it on the current query. We replace the deterministic tail with Poisson sampling at known inclusion probabilities, which makes the eviction error identifiable and turns a survey-sampling variance estimator over the retained set into a per-step error certificate at one extra scalar per retained token. On a thirty-turn assistant compressed to a 10\% cache budget, the certificate-gated system answers 0.97 of recall questions against 0.09 for top-$k$, and for facts stated 26 to 30 turns earlier it recalls 97\% against 2\%. We prove that no estimator computable from the information a deterministic scheme retains is consistent for its own eviction error: evicted values can be altered so that everything retained is unchanged while the true attention-output error grows without bound. Under the Poisson design the certificate covers the realized attention error in 96.9--97.7\% of 12{,}096 replay cells and in 98.1--99.7\% on twelve further architectures. Randomization buys attribution, not prediction: a pre-registered study on LongBench at 6k and 16k tokens (about 74{,}000 generations) finds question-aware eviction at 25--50\% budgets nearly free and output log-probability the better failure predictor, while the certificate answers the question confidence cannot, separating eviction-induced from inherent failures at AUC 0.65--0.75 against 0.47--0.54, and schedules recomputation at 1.7--1.8 times the gain of random gating. On real long-term conversations the gated system returns the full-cache score inside the heavy-damage regime, and the rule that triggers it is the same across five model families.