AI 中文总结
研究针对多轮大语言模型服务中键值缓存重用问题,提出HyMCache框架集成CXL混合内存。利用多轮缓存访问特性优化DRAM管理,通过请求级前缀预取等技术,在真实原型上评估,相比其他方案有性能优势且节省DRAM。
AI 中文摘要
长上下文、多轮和智能体大语言模型工作负载越来越多地重用先前处理的上下文,使键值缓存重用对于减少冗余计算至关重要。然而,这种重用将瓶颈转移到了在集群规模存储和提供可重用键值状态的内存层。GPU HBM和主机DRAM成本过高,无法扩展到TB级共享上下文容量,促使使用低成本、高容量介质构建远程层。本文提出了HyMCache,一种用于多轮大语言模型服务的集成CXL混合内存(CXL-HM)的键值缓存框架。CXL-HM在CXL接口后将少量设备内DRAM与大容量SSD支持的容量相结合。通过利用多轮键值缓存访问的读主导、可预测和仅追加的特性,HyMCache重新思考CXL-HM内的DRAM管理,以有效支持TB级SSD支持的键值重用。它使用请求级前缀预取和机会性写缓冲在设备DRAM中暂存对延迟至关重要的读取,以SSD级成本实现DRAM规模的键值缓存效率。我们在单聚合器和PD分解服务配置下的真实CXL-HM原型上评估了HyMCache。在相同的DRAM预算下,HyMCache在单节点服务中比本地LMCache性能高3.0倍,在PD分解服务中高1.45倍。与1TB分布式DRAM Mooncake相比,HyMCache性能低约30%,但使用的DRAM少16倍。
英文摘要
Long-context, multi-turn, and agentic LLM workloads increasingly reuse previously processed context, making KV-cache reuse essential for reducing redundant computation. However, this reuse shifts the bottleneck to the memory tier that stores and serves reusable KV states at cluster scale. GPU HBM and host DRAM are too costly to scale to TB-scale shared context capacity, motivating remote tiers built from lower-cost, higher-capacity media. This paper presents HyMCache, a CXL memory rack for multi-turn LLM serving. We build the memory rack using cost-efficient CXL-hybrid memory (CXL-HM), which combines a small amount of in-device DRAM with large SSD-backed capacity behind a CXL interface. By exploiting the read-dominant, predictable, and append-only nature of multi-turn KV-cache access, HyMCache rethinks DRAM management within CXL-HM to efficiently support TB-scale SSD-backed KV reuse. It uses request-level prefix prefetching and opportunistic write buffering to stage latency-critical reads in device DRAM, enabling DRAM-scale KV-cache efficiency at SSD-level cost. We evaluate HyMCache on a real CXL-HM prototype under both single-aggregator and PD-disaggregated serving configurations. Under the same DRAM budget, HyMCache outperforms local LMCache by 3.0x in single-node serving and 1.45x in PD-disaggregated serving. Compared with 1 TB distributed-DRAM Mooncake, HyMCache incurs about 30% lower performance but uses 16x less DRAM.