arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08684cs.LGcs.AI

RippleKV:基于扰动传播的跨层KV缓存分配

RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

Dongjie Xu, Kai Qian, Julius, Weijie Shi, Yuxuan Sun, Minghua Tang, Fenglei Jin, Hanchi Dong, Jiajie Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文LLM推理的KV缓存分配问题,提出RippleKV方法,通过层值缓存扰动敏感性分配预算,在LongBench上的匹配缓存预算下实现了最高平均性能。

中文摘要 AI 辅助

长上下文大语言模型(LLM)推理受限于KV缓存内存,如何在各层分配有限的缓存预算仍是难题。现有方法依赖层深度、注意力统计或表征变化等代理指标,但这些指标无法衡量各层扰动对输出的传播影响,可能导致敏感层缓存分配不足、耐受层分配过剩。为解决该问题,本文提出RippleKV,通过估计各层值缓存扰动对最终预测分布的影响来分配跨层缓存。RippleKV向各层值缓存独立注入归一化自适应扰动,在小型校准集上测量模型输出的诱导KL散度,对这些响应取平均得到与层深度无单调关系的模型专属敏感性分布。RippleKV通过归一化敏感性分数并应用指数映射,将敏感性分布转换为层预算乘数;比率参数控制敏感层与耐受层的分配差异,最终归一化操作则保留KV缓存总预算。在LongBench上的实验表明,在匹配的缓存预算下,RippleKV在所评估的KV缓存压缩方法中实现了最高的平均性能。

英文摘要

Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as layer depth, attention statistics, or representation change. These proxies do not measure how perturbations at each layer propagate to the output and may therefore cause sensitive layers to be underallocated while tolerant layers are overallocated. To address this issue, we propose RippleKV, which allocates cache across layers by estimating how perturbations to each layer's value cache affect the final predictive distribution. RippleKV independently injects norm-adaptive perturbations into each layer's value cache and measures the induced KL divergence at the model output over a small calibration set. Averaging these responses yields a sensitivity profile specific to the model that need not vary monotonically with depth. RippleKV then converts the sensitivity profile into layer budget multipliers by normalizing the sensitivity scores and applying an exponential mapping. A ratio parameter controls the allocation disparity between sensitive and tolerant layers, while a final normalization preserves the KV cache budget. Experiments on LongBench demonstrate that RippleKV achieves the highest average performance among the evaluated KV cache compression methods under matched cache budgets.

↑