自剪枝Transformer:基于通用注意力的极致KV缓存压缩
A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
浏览论文内容
中文总结 AI 辅助
提出通用注意力机制,通过复合衰减实现自适应KV缓存剪枝,在保持性能的同时达到10倍压缩,并在长上下文下实现25倍压缩。
中文摘要 AI 辅助
现代大语言模型的KV缓存规模庞大,成为高效部署的障碍。近期研究尝试用基于衰减的机制替代注意力层中的RoPE位置嵌入,并利用该机制在推理时对KV缓存进行剪枝。然而,这些衰减函数表达能力有限,在实践中退化为类似滑动窗口的驱逐模式。在本工作中,我们提出一个统一框架,用于设计互补且新颖的衰减机制,在保留高表达力RoPE嵌入和Softmax注意力的同时,捕获复杂的键统计信息和交互。由此产生的通用注意力(Universal Attention)是一种高表达力且可端到端训练的架构,其复合衰减机制作为自然的自适应剪枝准则,移除对注意力计算贡献最小的令牌。实验表明,通用注意力在自然语言和合成任务数据上实现了最先进的10倍压缩,同时在下游性能上优于最先进的基线和未剪枝的基准模型。此外,它在16k长度下展现出卓越的长上下文泛化能力,实现了前所未有的25倍压缩。
英文摘要
The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit{adaptive}$ pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $\textit{improving}$ downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented $25\times$ compression at length 16k.
发表机构
- IBM Research(IBM研究院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。