arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27981cs.CLcs.LG

风险控制的KV缓存驱逐:从内存预算到风险目标

Risk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targets

Beomgu Kang, SoJin Yun, Hojoon Kim, Hyunseok Seo

首次发表
浏览论文内容

中文总结 AI 辅助

针对KV缓存驱逐中平均损失掩盖个别请求效用显著下降的问题,提出将驱逐重构为部署风险控制问题,采用与压缩器无关的事后认证程序,在有限样本保证下选择保留策略,并在无认证策略时回退全KV,从而将可靠性要求转化为内存运行点。

中文摘要 AI 辅助

KV缓存驱逐通常通过平均质量-内存权衡来评估,然而,一个很小的平均损失可能掩盖那些效用显著下降的请求。我们将驱逐重新表述为一个部署风险控制问题:当驱逐相对于同一请求上的全KV推理,将任务效用降低超过部署指定的容差时,即发生实质性退化,而部署风险则是此类事件在总体中的发生频率。给定一个指定目标风险水平和置信度要求的可靠性契约,我们采用一种与压缩器无关的事后认证程序,从校准数据中选择一个保留策略,并具有有限样本保证;当没有压缩策略通过认证时,则回退到全KV。在多种驱逐方法、Llama和Mistral模型以及LongBench和RULER-32K基准上,相同的契约支持显著不同的驱逐水平:在Llama上,它认证了SnapKV在LongBench上的75%保留率,但在RULER-32K上,没有任何测试的压缩策略通过认证,从而触发全KV回退。经验退化率低于5%目标的策略仍可能无法通过有限样本认证;在Llama LongBench上,经验阈值选择会选出未认证的策略,这些策略在固定预算方法下保留的缓存少5-10个百分点。所提出的框架将部署级别的可靠性要求转化为KV内存运行点。

英文摘要

KV-cache eviction is typically evaluated through average quality-memory trade-offs, yet a small average loss can hide requests whose utility degrades materially. We reformulate eviction as a deployment risk-control problem: a material degradation occurs when eviction lowers task utility by more than a deployment-specified tolerance relative to full-KV inference on the same request, and deployment risk is the population frequency of such events. Given a reliability contract specifying a target risk level and confidence requirement, we use a compressor-agnostic post-hoc certification procedure to select a retention policy from calibration data with a finite-sample guarantee, falling back to full KV when no compressed policy is certified. Across multiple eviction methods, Llama and Mistral models, and LongBench and RULER-32K, the same contract supports substantially different levels of eviction: on Llama, it certifies SnapKV at 75% retention on LongBench but no tested compressed policy on RULER-32K, triggering full-KV fallback. Policies with empirical degradation rates below the 5% target can still fail finite-sample certification; on Llama LongBench, empirical thresholding selects uncertified policies that retain 5-10 percentage points less cache across fixed-budget methods. The proposed framework converts a deployment-level reliability requirement into a KV-memory operating point.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑