arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02029cs.AI

HeadWiseKV:面向混合长上下文语言模型的按头预算键值缓存驻留方案

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

  • Nanjing University of Posts and Telecommunications(南京邮电大学)
  • Tylogi AI Lab / TAIL(Tylogi AI实验室/TAIL)
  • Anhui University of Technology(安徽工业大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • School of Information Science and Engineering, Southeast University(东南大学信息科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu

AI总结:

HeadWiseKV是一种无需训练的框架,通过SeqCalib算法分配静态多级历史窗口,在保留混合长上下文语言模型核心路径的同时,降低KV缓存内存占用并扩展支持的上下文长度。

AI中文摘要:

长上下文语言模型在解码过程中会持续保留键值(KV)缓存,该缓存会消耗大量GPU内存并降低生成吞吐量。混合语言模型中仍存在这一瓶颈,因为其残差全局注意力层会主导依赖上下文的缓存需求。本文研究在总KV驻留预算下如何分配该状态,提出HeadWiseKV——一种无需训练的框架,用于压缩混合语言模型的残差全局KV缓存,同时保留其原生的局部、循环和线性路径。该框架为每个物理KV头分配静态多级历史窗口,使缓存需求在服务前即可预测。我们将该分配问题形式化为受限操作率失真问题,并提出SeqCalib作为HeadWiseKV的核心策略生成算法。SeqCalib按执行顺序处理各层,且每个决策都基于部署时使用的下层策略,从而考虑了跨深度的交互。分组缓存运行时将选定策略实现为实际的按头KV驻留,而非完整缓存上的掩码。我们在四个混合长上下文模型上评估下游质量,并在Qwen3.6-27B上研究物理驻留和服务行为。HeadWiseKV在所有评估模型上保持了接近全KV缓存的RULER和LoCoMo质量;在固定模型系统研究中,其在112K上下文长度下将采样峰值设备内存降低8.59%,并将已验证的最大成功上下文长度从114K扩展至161K。

英文摘要:

Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths. It assigns each physical KV head a static, multilevel history window, making cache demand predictable before serving. We formulate this allocation as a restricted operational rate--distortion problem and propose SeqCalib as the core policy-generation algorithm in HeadWiseKV. SeqCalib processes layers in execution order and conditions each decision on the lower-layer policy used at deployment, thereby accounting for interactions across depth. A grouped-cache runtime materializes the selected policy as actual per-head KV residency rather than a mask over a full cache. We evaluate downstream quality across four hybrid long-context models and study physical residency and serving behavior on Qwen3.6-27B. HeadWiseKV retains near-Full-KV RULER and LoCoMo quality across the evaluated models. In the fixed-model systems study, it reduces sampled peak device memory by 8.59\% at a 112K context length and extends the largest verified successful context from 114K to 161K.

补充信息

↑