arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08131cs.CL

Jacap:基于雅可比非线性信息容量保持的稳健KV缓存淘汰

Jacap: Robust KV Cache Eviction via Jacobian-Based Nonlinear Information Capacity Preservation

Jiaming Yang, Chenwei Tang, Liangli Zhen, Chenyang Zhang, Jiancheng Lv

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Jacap,一种基于雅可比信息容量的KV缓存淘汰方法,通过非线性信道建模和容量感知选择,在高压缩率下实现优越性能。

中文摘要 AI 辅助

键值(KV)缓存淘汰对于扩展大型语言模型中的长上下文推理至关重要。然而,现有策略主要依赖经验启发式方法,缺乏在固有非线性softmax注意力机制下对令牌效用的严格刻画。在本工作中,我们从局部信息几何的角度重新思考KV缓存淘汰,将注意力过程建模为非线性高斯通信信道。通过对注意力映射进行一阶泰勒展开,我们推导出雅可比信息容量,这是一个新颖的目标,明确捕获了查询相关性、softmax敏感性和结构多样性。在此理论的指导下,我们提出了Jacap,一种容量感知的淘汰方法,利用softmax感知的重要性加权和统计杠杆分数进行子集选择。跨多种架构和基准的大量实验表明,Jacap在大多数场景中表现出优越的性能,尤其是在高压缩率情况下。

英文摘要

Key-value (KV) cache eviction is essential for scaling long-context inference in Large Language Models. However, existing policies predominantly rely on empirical heuristics, lacking a rigorous characterization of token utility under the inherently nonlinear softmax attention mechanism. In this work, we rethink KV cache eviction through the lens of local information geometry, modeling the attention process as a nonlinear Gaussian communication channel. By performing a first-order Taylor expansion of the attention mapping, we derive the Jacobian Information Capacity, a novel objective that explicitly captures query relevance, softmax sensitivity, and structural diversity. Guided by this theory, we introduce Jacap, a capacity-aware eviction method that utilizes softmax-aware importance weighting and statistical leverage scores for subset selection. Extensive experiments across diverse architectures and benchmarks demonstrate that \textsc{Jacap} delivers superior performance in most scenarios, particularly in high-compression regimes.

发表机构

  • College of Computer Science, Sichuan University(四川大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑