arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26300cs.LGcs.AI

CompKV:面向长上下文LLM推理的补偿感知KV选择

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu

首次发表
浏览论文内容

中文总结 AI 辅助

CompKV提出首个补偿感知的稀疏注意力框架,通过块级统计优化令牌选择以最小化补偿误差,在长上下文推理中实现最佳性能与高达6.85倍加速。

中文摘要 AI 辅助

尽管大型语言模型(LLMs)性能强大,但在长上下文推理过程中,其性能受限于KV缓存内存流量。稀疏注意力被广泛用于加速LLM推理,通过计算选定令牌子集上的精确注意力来实现。为了恢复被排除在精确注意力之外的令牌的贡献,近期方法对省略的注意力尾部应用粗粒度补偿。然而,现有方法通常基于注意力质量选择令牌,然后才补偿未选中的令牌。这种解耦设计忽视了它们之间的相互作用:选择应优先考虑那些如果被省略会留下最大补偿误差的令牌。为解决这一局限,我们引入了CompKV,这是首个补偿感知的稀疏注意力框架,它将令牌划分为块,并显式优化选择以适配下游补偿机制。我们的理论分析表明,块级均值补偿留下的残差由块注意力质量和块内logit变化共同决定。我们利用紧凑的块级统计量近似该残差,从而得到一个可部署的选择标准。我们进一步开发了一种高效的异步实现。在RULER和LongBench-Pro上的实验表明,CompKV在评估的稀疏基线中表现最佳,同时相比全注意力实现了高达6.85倍的自我注意力加速。

英文摘要

Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens. To recover the contribution of tokens excluded from exact attention, recent methods apply coarse-grained compensation to the omitted attention tail. However, existing methods typically select tokens based on attention mass and only then compensate for the unselected tokens. This decoupled design overlooks their interaction: selection should prioritize tokens that would leave the largest compensation error if omitted. To address this limitation, we introduce CompKV, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism. Our theoretical analysis shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit variation. We approximate this residual using compact block-level statistics, yielding a deployable selection criterion. We further develop an efficient asynchronous implementation. Experiments on RULER and LongBench-Pro show that CompKV performs best among the evaluated sparse baselines while delivering up to a $6.85\times$ self-attention speedup over full attention.

发表机构

  • Tsinghua University(清华大学)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑