arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17715cs.CL

C$^2$KV:用于高效大语言模型推理的压缩可组合键值缓存重用

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

发表机构阿里巴巴集团
查看机构详情
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Chuheng Du, Junyi Chen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chaoyue Niu, Shengzhong Liu, Guihai Chen, Fan Wu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对长上下文LLM推理成本高问题,提出C$^2$KV框架,联合优化KV提取与推理拼接,学习可组合压缩的KV缓存流形,引入轻量级边车提取器及协同训练策略,显著降低缓存成本,实现推理加速并保持生成质量。

中文摘要 AI 辅助

长上下文推理是现代大语言模型(LLM)应用(如检索增强生成和多文档推理)的核心。为减轻不断增长的推理成本,近期工作探索了键值(KV)缓存重用以减少冗余预填充计算。但现有方法主要关注计算节省,忽视了长上下文LLM服务中的关键瓶颈:存储和访问大型KV缓存的成本。虽然KV压缩似乎是自然补充,但将其与非前缀KV重用简单结合常导致严重的精度下降。在这项工作中,我们提出C$^2$KV,一个用于非前缀KV重用的统一框架,联合优化KV提取和推理时的拼接。C$^2$KV学习一个可组合且压缩的KV缓存流形,明确设计为与位置无关。我们的方法引入了一个轻量级的带有可学习压缩令牌和结构化注意力流的边车提取器,实现模块化的KV表示,可灵活重用和拼接而无需修改冻结的基础模型。我们还采用了压缩 - 拼接协同训练策略,使提取时的表示与其下游重用行为对齐。跨多个长上下文基准和模型系列的广泛实验表明,C$^2$KV显著降低了KV缓存存储和传输成本,在长上下文下实现高达17倍的推理加速,同时保持生成质量。

英文摘要

Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value (KV) cache reuse to reduce redundant prefill computation. However, existing reuse methods primarily focus on computation savings and overlook a critical bottleneck in long-context LLM serving: the cost of storing and accessing large KV caches. While KV compression appears to be a natural complement, naively combining compression with non-prefix KV reuse often leads to severe accuracy degradation. In this work, we propose C$^2$KV, a unified framework for non-prefix KV reuse that jointly optimizes KV extraction and inference-time concatenation. C$^2$KV learns a composable and compressed KV cache manifold that is explicitly designed to be position-agnostic. Our approach introduces a lightweight sidecar Extractor with learnable compression tokens and a structured attention flow, enabling modular KV representations that can be flexibly reused and concatenated without modifying the frozen base model. We further employ a compression-concatenation co-training strategy to align extraction-time representations with their downstream reuse behavior. Extensive experiments across multiple long-context benchmarks and model families demonstrate that C$^2$KV significantly reduces KV cache storage and transfer costs, achieving up to 17$\times$ inference speedup under long contexts, while preserving generation quality.

补充信息

↑