发表机构
USC; KAIST; Amazon AGI(南加州大学; 韩国科学技术院; 亚马逊AGI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLMs长上下文推理的KV缓存内存瓶颈,提出SGD-KV框架,通过分块摘要诊断任务识别分层信息聚合的注意力头,在Qwen2.5模型上实现1M token上下文的SOTA性能,同时降低75%内存使用,达成更优效率-精度权衡。
AI 中文摘要
大语言模型(LLMs)在长上下文推理中面临严重的内存瓶颈,原因是键值(KV)缓存的规模呈线性增长。现有的KV缓存压缩技术通常依赖简单的启发式方法,忽略了不同注意力头的不同功能角色。我们提出SGD-KV(Summarization-Guided KV Cache Compression,即摘要引导的KV缓存压缩),这是一种感知注意力头的框架,利用新颖的分块摘要诊断任务,系统识别并优先处理专门用于分层信息聚合的注意力头。在Qwen2.5-7B-1M和Qwen3-32B模型上,针对多种长上下文基准的实验表明,SGD-KV在处理长达100万token的上下文时达到了最先进的性能,同时将KV缓存的内存使用量降低了高达75%。我们的研究结果表明,基于注意力头的摘要得分分布策略性分配KV缓存预算,可为长上下文推理实现更优的效率-准确性权衡。
英文摘要
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We present SGD-KV (Summarization-Guided KV Cache Compression), a head-aware framework that leverages a novel chunk-summarization diagnostic task to systematically identify and prioritize attention heads specialized in hierarchical information aggregation. Experiments on Qwen2.5-7B-1M and Qwen3-32B across diverse long-context benchmarks demonstrate that SGD-KV achieves state-of-the-art performance with contexts up to 1M tokens, while reducing KV cache memory usage by up to 75%. Our findings show that strategically allocating the KV cache budget based on the summarization score distribution of attention heads yields a superior efficiency-accuracy trade-off for long-context inference.
CommentsAccepted in NeurIPS2026 Efficient Reasoning Workshop