arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Distance-KV:利用相对距离实现高效长上下文推理

Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference

Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan, Qianyi Cai, Xiaojun Quan, Xiangdong Su

arXiv 2609.32663首次发表:更新:

发表机构

Inner Mongolia University; Shenzhen Loop Area Institute; Sun Yat-sen University; Kyoto University; The Hong Kong University of Science and Technology (Guangzhou)(内蒙古大学; 深圳河套学院; 中山大学; 京都大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Distance-KV通过离线学习层、注意力头和相对距离上的静态KV保留模式,在不进行在线重要性评分的情况下压缩KV缓存,在长上下文基准上超越现有压缩方法,显著降低内存并加速解码。

AI 中文摘要

大语言模型推理的内存占用和解码延迟随上下文长度快速增长。为降低这些成本,键值(KV)缓存压缩方法根据令牌重要性或各注意力头间注意力模式的差异,选择性地保留缓存状态。然而,我们发现,即使在同一注意力头内,检索能力也会随相对距离发生显著变化。为利用这一结构,我们提出了Distance-KV,该方法在层、注意力头和相对距离的联合空间上学习一个静态的KV保留模式。该模式在冻结语言模型的情况下离线学习,并跨输入复用,以在不进行在线重要性评分的情况下剪枝和压缩KV缓存。在三个骨干模型和四个长上下文基准上,Distance-KV在竞争的KV缓存压缩方法中始终取得最佳整体性能,在128K的RULER上比最强压缩基线高出最多9.3个点。在128K的Llama-3.1-8B-Instruct上,Distance-KV将KV缓存内存减少65.4%,并相对于Dense实现了1.66倍的解码加速。综合这些结果,相对距离被确定为理解大语言模型如何在长上下文中检索信息以及设计更高效推理方法的重要结构维度。

英文摘要

The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies substantially with relative distance, even within the same attention head. To exploit this structure, we introduce Distance-KV, which learns a static KV retention pattern over the joint space of layers, attention heads, and relative distances. The pattern is learned offline with the language model frozen and reused across inputs to prune and compact the KV cache without online importance scoring. Across three backbone models and four long-context benchmarks, Distance-KV consistently achieves the best overall performance among competing KV cache compression methods, exceeding the strongest compression baseline by up to 9.3 points on RULER at 128K. On Llama-3.1-8B-Instruct at 128K, Distance-KV reduces KV cache memory by 65.4% and achieves a $1.66\times$ decoding speedup relative to Dense. Together, these results identify relative distance as an important structural dimension for understanding how LLMs retrieve information over long contexts and for designing more efficient inference methods.

Comments19 pages, 7 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑