Recency/Frequency Adaptive KV Caching for Large Language Model Serving
面向大语言模型服务的近期/频率自适应KV缓存
机构 * Department of Computer Science, Johns Hopkins University, Baltimore, USA(约翰霍普金斯大学计算机科学系) ; Argonne National Laboratory, Lemont, USA(阿贡国家实验室) ; Parasail, Inc., San Mateo, USA(Parasail公司)
专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG;foundation model(comments)
AI总结 针对LRU策略在多工作负载下缓存失效问题,提出自适应KV缓存,动态分配近期与高频KV块的空间,在合成文档问答中缓存命中率提升10.8%,首令牌时间降低12.6%。
Comments Accepted at the ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)