发表机构
University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过概率推理形式化KV驱逐问题,证明其计算难度,将其简化为可采样近似的期望估计,提出解码时校正方法,调整现有方法实现该校正,使结合该校正的KV驱逐在相同压缩预算下更鲁棒且性能有竞争力。
AI 中文摘要
KV(缓存)驱逐的前提和目标很简单:通过从KV缓存中驱逐部分条目,可在对质量影响可忽略不计的情况下实现更高的吞吐量,这在许多现有方法的实验中得到了验证,尽管大多数方法依赖于创造性的启发式方法来选择要丢弃的条目。尽管最近取得了一些进展,但KV驱逐问题在文献中仍未得到正式定义。本文旨在通过概率推理的视角对该问题进行适当的形式化,并揭示从该视角能获得的见解。具体而言,我们(1)形式化KV驱逐问题,遗憾的是,证明其具有计算难度;(2)表明通过将其框架化为概率问题,KV驱逐可简化为期望估计问题,该问题可通过采样进行近似;(3)表明通过这种概率解释,在解码过程中对被驱逐条目进行校正——这一此前被忽视的问题——变得可行;(4)揭示文献中的现有方法是零方差有偏估计器,可轻松调整以实现解码时校正。在实践中,我们表明,这种结合了解码时校正的概率版KV驱逐,与现有驱逐方法相比,对不同任务更具鲁棒性,且在相同压缩预算下可达到有竞争力的性能。
英文摘要
The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selecting which entries to drop. Despite recent advances, the problem of KV eviction has remained informal in the literature. This paper aims to properly formalize this problem through the lens of probabilistic reasoning and reveal what can be learned from this perspective. Concretely, we (1) formalize the problem of KV eviction and, unfortunately, prove that it is computationally hard, (2) show that by framing it probabilistically, KV eviction reduces to the problem of expectation estimation, which can be approximated through sampling, (3) show that through this probabilistic interpretation, correcting for evicted entries during decoding---a previously ignored problem---becomes feasible, and (4) reveal that existing methods in the literature are zero-variance biased estimators that can be easily adapted in order to enable decode time correction. In practice, we show that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.