arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10790cs.DCcs.LG

可组合CXL内存作为Kubernetes原生共享内存用于LLM服务

Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

  • Seagate Technology, LLC(希捷科技有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra

AI总结:

本文提出一种Kubernetes DRA驱动,使可组合CXL内存成为可调度资源,实现跨节点KV缓存共享,显著降低LLM服务TTFT,验证了内存解聚的可行性。

AI中文摘要:

我们提出了一种Kubernetes动态资源分配(DRA)驱动程序,使可组合CXL内存成为可调度的集群资源,并评估了由此产生的共享内存层在LLM服务中用于跨节点KV缓存复用。该驱动程序按需组合CXL区域,将其物化为每个参与主机上的DAX设备,并在单一容器设备接口(CDI)名称下注入到Pod中,使得不同节点上的Pod能够访问同一物理区域。一个针对vLLM/llm-d的共享内存连接器将该区域用作KV缓存层,并在共享介质内嵌入槽目录,从而无需外部元数据服务。在一个配备512 GiB CXL设备、运行Qwen2.5-7B-Instruct的双节点集群上,跨节点前缀复用将TTFT降低了5.5倍至36.6倍,外部命中率为95.4%至99.5%,而节点本地层(GPU前缀缓存、CPU-DRAM卸载)则退化为完全重新计算。共享差距(定义为跨节点与同节点复用之间的延迟比)为1%至4%,表明在我们的测试平台上,跨节点复用相比同节点复用仅产生极小的额外延迟。两个副本均运行完整引擎;该研究展示了内存解聚而非预填充/解码解聚。我们将其报告为可行性研究而非性能评估。

英文摘要:

We present a Kubernetes Dynamic Resource Allocation (DRA) driver that makes composable CXL memory a schedulable cluster resource, and evaluate the resulting shared-memory tier for cross-node KV-cache reuse in LLM serving. The driver composes CXL regions on demand, materializes them as DAX devices on each participating host, and injects them into pods under a single Container Device Interface (CDI) name so that pods on different nodes access the same physical region. A shared-memory connector for vLLM/llm-d uses that region as a KV-cache tier with a slot directory embedded inside the shared medium, which eliminates the need for an external metadata service. On a two-node cluster with a 512\,GiB CXL appliance and Qwen2.5-7B-Instruct, cross-node prefix reuse reduces TTFT by 5.5$\times$--36.6$\times$ at an external hit rate of 95.4--99.5\,\%, while node-local tiers (GPU prefix caching, CPU-DRAM offload) fall back to full recompute. The sharing gap, defined as the latency ratio between cross-node and same-node reuse, is 1--4\%, indicating that cross-node reuse incurs little additional latency relative to same-node reuse on our testbed. Both replicas run full engines; the study demonstrates memory disaggregation rather than prefill/decode disaggregation. We report this as a feasibility study rather than a performance evaluation.

补充信息

↑