arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CateKV:面向长上下文大语言模型推理加速的顺序一致性研究

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

Haoyun Jiang, Haolin Li, Jianwei Zhang, Fei Huang, Qiang Hu, Minmin Sun, Shuai Xiao, Yong Li, Junyang Lin, Jiangchao Yao

arXiv 2608.30295首次发表:更新:

发表机构

Shanghai Jiao Tong University; Alibaba Group; Fudan University(上海交通大学; 阿里巴巴集团; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对长上下文LLM推理的内存与延迟挑战,提出混合KV缓存方法CateKV,利用注意力头的顺序一致性特性减少冗余KV信息,在保持准确率的同时显著降低内存占用、提升解码与批量处理效率。

AI 中文摘要

大语言模型(LLMs)在处理长上下文任务时展现出强大能力,但处理这类长上下文因内存需求大、推理延迟高而颇具挑战。本研究发现部分注意力头的注意力模式存在顺序一致性,可通过基于变异系数的算法持续识别。受此启发,我们提出CateKV,这是一种混合KV缓存方法,仅为具有顺序一致性的注意力头保留关键token信息,从而减小KV缓存规模、降低计算开销,同时为自适应注意力头保留大部分KV对以确保高准确率。我们展示了该算法的独特特性及其与现有加速方法的扩展应用。在长上下文基准测试中的全面评估表明,在保持与全注意力相当的准确率的同时,CateKV在单样本输入下将内存使用减少多达2.72倍,解码速度提升2.18倍,在批量场景中吞吐量提升3.96倍。

英文摘要

Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm. Inspired by this observation, we propose CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead, while preserving the majority of KV pairs in adaptive heads to ensure high accuracy. We show the unique characteristics of our algorithm and its extension with existing acceleration methods. Comprehensive evaluations on long-context benchmarks show that, while maintaining accuracy comparable to full attention, CateKV reduces memory usage by up to $2.72\times$ and accelerates decoding by $2.18\times$ in single-sample inputs, and boosts throughput by $3.96\times$ in batch scenarios.

CommentsPublished at ICML 2025

Journal refProceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267:27569-27585, 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑