发表机构
Infinigence-AI; Tsinghua University; Shanghai Jiao Tong University; Peking University(智谱AI; 清华大学; 上海交通大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PDD提出三层跨数据中心预填充-解码分离架构,通过RDMA重叠传输与解码移交机制,在异构GPU上实现高达37.5%的效益成本比提升。
AI 中文摘要
通过广域以太网互连地理上分散的集群,为大型语言模型(LLM)推理提供了一种可扩展且成本效益高的替代方案,以替代专用的数据中心内异构集群。然而,这种跨数据中心分离在集群之间施加了沉重的KV缓存传输,导致显著的首令牌时间(TTFT)延迟,这对于以长上下文、高缓存命中率和短输出为特征的智能体工作负载尤其不利。我们提出了PDD,一种基于预填充-解码(PD)分离构建的三层分离架构,由预填充(Prefill)、中继解码(RLD)和主解码(MD)实例组成。在集群A上,预填充实例执行预填充计算,而RLD实例通过高速RDMA立即接收KV缓存并开始解码,从而与通过基于TCP的以太网向集群B传输KV重叠。集群B上的MD实例接收KV缓存以及RLD产生的令牌,然后解码从RLD无缝移交给MD以完成。为了最大化整体效率,PDD采用了三种核心机制:解码端RadixCache以缓解带宽瓶颈,扩展-解码移交机制以实现RLD和MD之间的平滑控制迁移,以及多阶段流水线编排以管理高并发和长期服务下的复杂层间依赖。我们进一步设计了一种低成本、细粒度的异构部署方案,以边际成本最大化延迟掩蔽效率。与数据中心内同构PD基线相比,PDD对计算密集型H100和内存带宽优化型H200的跨数据中心映射,在符合SLA的良好吞吐量方面实现了高达37.5%更高的效益成本比(BCR)。
英文摘要
Interconnecting geographically dispersed clusters over wide-area Ethernet provides a scalable and cost-effective alternative to dedicated intra-datacenter heterogeneous clusters for Large Language Model (LLM) inference. However, this cross-datacenter disaggregation imposes heavy KV-cache transfers between clusters, causing substantial Time-to-First-Token (TTFT) latency, which is especially detrimental for agentic workloads characterized by long contexts, high cache hit rates, and short outputs. We propose PDD, a three-tier disaggregation architecture built upon prefill-decode (PD) disaggregation, consisting of Prefill, RelayDecode (RLD), and MainDecode (MD) instances. On Cluster A, Prefill instances perform the prefill computation, while RLD instances immediately receive the KV cache via high-speed RDMA and begin decoding, thereby overlapping the KV transfer to Cluster B over TCP-based Ethernet. MD instances on Cluster B receive the KV cache along with the tokens produced by RLD, and decoding is then seamlessly handed off from RLD to MD for completion. To maximize overall efficiency, PDD employs three core mechanisms: Decode-side RadixCache to alleviate bandwidth bottlenecks, an Extend-Decode Handoff mechanism for smooth control migration between RLD and MD, and multi-stage pipeline orchestration to manage complex inter-tier dependencies under high concurrency and long-term serving. We further design a low-cost, fine-grained heterogeneous deployment scheme that maximizes latency-masking efficiency at marginal cost. Compared to the intra-DC homogeneous PD baseline, PDD's cross-datacenter mapping of compute-intensive H100s and memory-bandwidth-optimized H200s achieves a Benefit-Cost Ratio (BCR) up to 37.5% higher in SLA-compliant goodput.