Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
迈向高效大语言模型服务:关于系统感知键值缓存优化的综述
机构 * School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院) ; School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
AI总结 该综述聚焦大语言模型服务中系统感知的键值缓存优化,从执行与调度、放置与迁移、表示与保留三个维度回顾相关工作,分析跨行为协同设计及行为与目标联系,为理解和创新键值缓存设计提供基础。
Comments Accepted to ACL 2026 as a Findings paper
Journal ref Findings of the Association for Computational Linguistics: ACL 2026 (pp. 38450-38476)