CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration
CateKV:面向长上下文大语言模型推理加速的顺序一致性研究
机构 * Shanghai Jiao Tong University(上海交通大学) ; Alibaba Group(阿里巴巴集团) ; Fudan University(复旦大学)
AI总结 本研究针对长上下文LLM推理的内存与延迟挑战,提出混合KV缓存方法CateKV,利用注意力头的顺序一致性特性减少冗余KV信息,在保持准确率的同时显著降低内存占用、提升解码与批量处理效率。
Comments Published at ICML 2025
Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267:27569-27585, 2025