面向拆分式GPU推理的拓扑感知数据移动
Topology-Aware Data Movement for Disaggregated GPU Inference
AI总结:
针对拆分式GPU推理中现有系统忽略GPU间带宽差异的问题,本文设计拓扑感知传输编排器,通过三项机制实现3至18倍传输延迟降低。
AI中文摘要:
拆分式大语言模型(LLM)推理会产生现有系统无法正确解决的数据中心网络问题:当预填充(prefill)和解码(decode)在不同GPU池运行时,必须在两者间传输KV缓存。对于70B模型,单请求KV缓存达2.6GB,生产规模下总聚合带宽超100GB/s。然而DistServe、Splitwise、Mooncake均采用统一的RDMA传输,忽略了GPU间带宽因物理关系差异达72倍:同一域内通过NVLink达900GB/s,跨节点通过InfiniBand达50GB/s,跨数据中心通过TCP达12.5GB/s。本文设计了拓扑感知传输编排器,其启动时会发现互连层级并为每次传输选择最优传输方式,包含三项协同机制:(1)逐流水线层传输,将传输与正在进行的预填充重叠,可隐藏60%至85%的延迟于计算之后;(2)针对混合专家(Mixture-of-Experts)模型的NVLink域感知放置,协同优化专家调度与KV缓存局部性;(3)采用CXL 3.0内存扩展器作为共享溢出层,提供6倍容量且延迟比NVMe低86倍。完整评估需具备异构互连与CXL 3.0硬件的多节点集群,这类资源超出学术机构能力且未在GPU云提供,本文提出了分析带宽模型、组件实现及三种架构下的预测分析,显示其相比统一RDMA可降低3至18倍的传输延迟。
英文摘要:
Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. For a 70B model this is 1.3 GB per request, exceeding 100 GB/s aggregate at production scale. Yet DistServe, Splitwise, and Mooncake all use uniform RDMA, ignoring that bandwidth between two GPUs varies by 72x depending on their physical relationship: 900 GB/s via NVLink 4.0 within a domain (1.8 TB/s on NVLink 5, widening the gap to 144x), 50 GB/s via InfiniBand across nodes, 12.5 GB/s via TCP across data centers. We design a topology-aware transfer orchestrator that discovers interconnect hierarchy at startup and selects optimal transport per transfer. Three mechanisms work together: (1) pipelined layer-by-layer transfer that overlaps transmission with ongoing prefill, hiding 76 to 100 percent of transfer latency behind computation depending on transport, with NVLink and PCIe transfers hidden entirely; (2) NVLink domain-aware placement for Mixture-of-Experts models that co-optimizes expert dispatch with KV cache locality; and (3) CXL 3.0 memory expanders as a shared overflow tier providing 6x capacity at 86x lower latency than NVMe. Full evaluation requires multi-node clusters with heterogeneous interconnects and CXL 3.0 hardware that is beyond academic resources and not yet available in GPU clouds. We present analytical bandwidth models, component implementations, and projected analysis across three architectures showing 3 to 18x transfer latency reduction over uniform RDMA.