发表机构
Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生成式推荐服务,提出写感知的LRU-K KV缓存策略,通过准入控制减少写入流量,使高带宽闪存吞吐量提升3.8-4.7倍,寿命从一年延至六年以上。
AI 中文摘要
生成式推荐(GR)系统日益利用用户级KV缓存重用,以避免重新计算冗长的用户历史记录。然而,不断增长的KV缓存容量和带宽需求给内存系统带来了新的挑战。高带宽闪存(HBF)通过提供远高于HBM的容量,同时接近HBM级别的读带宽,提供了一个有前景的解决方案,从而支持更大规模的KV缓存保留并提高服务吞吐量。然而,传统的最近最少使用(LRU)KV缓存管理将KV缓存写入与缓存未命中紧密耦合,产生过多的写入流量,迅速耗尽闪存寿命。在本工作中,我们评估了一种基于准入控制的LRU-K的写感知KV缓存策略,用于基于HBF的GR服务。通过在缓存准入前过滤低重用用户,LRU-K将KV缓存写入与未命中解耦,显著减少了不必要的写入。我们开发了一个分析模型来表征GR服务性能、KV缓存写入流量和HBF寿命,并在多种内存系统和GR工作负载下评估性能。我们的结果表明,基于HBF的系统比仅HBM的系统实现了3.8至4.7倍的吞吐量提升。此外,LRU-K将HBF寿命从传统LRU下的大约一年延长至K=10适中参数下的六年以上,同时保持相当或甚至略有提升的吞吐量。这些结果凸显了写感知KV缓存策略对于可持续的基于HBF的GR服务的重要性。
英文摘要
Generative recommendation (GR) systems increasingly leverage user-level KV cache reuse to avoid recomputing long user histories. However, the growing KV cache capacity and bandwidth requirements introduce new challenges for memory system. High-Bandwidth Flash (HBF) provides a promising solution by offering substantially higher capacity than HBM while approaching HBM-class read bandwidth, enabling larger scale KV cache retention and improved serving throughput. Yet conventional Least-Recently-Used (LRU) KV cache management tightly couples KV cache writes with cache misses, generating excessive write traffic that rapidly exhausts flash endurance. In this work, we evaluate a write-aware KV cache policy based on admission-controlled LRU-K for HBF-based GR serving. By filtering low-reuse users before cache admission, LRU-K decouples KV cache writes from misses and significantly reduces unnecessary writes. We develop an analytical model to characterize GR serving performance, KV cache write traffic, and HBF lifetime, and evaluate performance across diverse memory systems and GR workloads. Our results show that HBF-based systems achieve 3.8 to 4.7 times higher throughput than HBM-only systems. Moreover, LRU-K extends HBF lifetime from about one year under conventional LRU to over six years with a moderate K=10, while maintaining comparable or even slightly improved throughput. These results highlight the importance of write aware KV cache policy for sustainable HBF-based GR serving.
Comments5 pages, 7 figures