arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21079cs.DCcs.SYeess.SY

DLB:面向生成式AI推理的大规模分布式负载均衡

DLB: Distributed Load Balancing at Scale for Generative AI Inference

Santiago R. Balseiro, Bartek Wydrowski, Sameer Agarwal, David Applegate, Aaron Archer, Soheil Hassas Yeganeh, Alex Iriza, Bobby Kleinberg, Balasubramanian Sivan… 展开作者

Santiago R. Balseiro, Bartek Wydrowski, Sameer Agarwal, David Applegate, Aaron Archer, Soheil Hassas Yeganeh, Alex Iriza, Bobby Kleinberg, Balasubramanian Sivan, Pranav Vaish, Oscar Zegarra, Wenxin Zhang, Vahab Mirrokni, Amin Vahdat

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DLB分布式负载均衡系统,通过点对点探测和延迟模型学习,在Google大规模生成式AI推理中实现中位延迟降低17%、p95尾部延迟降低13%。

中文摘要 AI 辅助

现代数据中心对GPU和TPU等稀缺且昂贵的加速器的依赖,给后端基础设施带来了前所未有的压力。对于以异构服务时间和复杂多阶段处理为特征的工作负载(如生成式AI),传统的负载均衡技术往往力不从心,严重依赖成本高昂的过度配置来维持服务水平目标。本文介绍了DLB(分布式负载均衡器),这是一个新颖的系统,旨在为大规模异构工作负载最小化端到端用户延迟。DLB采用可扩展的分布式设计,通过点对点探测来实时了解跨大规模地理分布式基础设施的服务器容量。该系统持续学习延迟模型,以估计路由决策对延迟的影响,从而有效管理异构硬件和多样化的模型架构。我们对路由算法进行了新颖的理论分析,确立了其随时间推移的稳定性和全局性能保证。我们还通过大量模拟评估了DLB,结果显示与最先进的负载均衡算法相比,DLB带来了显著提升。最后,在DLB于Google部署22个月后(期间它支持了数千种不同机器学习模型的大规模生成式AI推理,每秒处理数百万个请求),我们详细介绍了系统在生产环境中的设计选择和实践经验。对生产迁移的分析表明,与旧基线相比,DLB在延迟方面取得了统计上显著的降低,其中中位延迟降低了17%,p95尾部延迟降低了13%。

英文摘要

The reliance on scarce and expensive accelerators such as GPUs and TPUs in modern datacenters places unprecedented demands on backend infrastructure. For workloads characterized by heterogeneous service times and complex multi-stage processing, such as Generative AI, conventional load balancing techniques are often inadequate, relying heavily on costly overprovisioning to maintain service level objectives. This paper introduces DLB, the Distributed Load Balancer, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads. DLB employs a scalable, distributed design with peer-to-peer probing to maintain real-time visibility into server capacity across large-scale, geographically distributed infrastructure. The system continuously learns latency models to estimate the latency impact of routing decisions, allowing it to effectively manage heterogeneous hardware and diverse model architectures. We provide a novel theoretical analysis of our routing algorithms that establishes their stability and global performance guarantees over time. We also evaluate DLB through extensive simulations, which show substantial gains compared to state-of-the-art load balancing algorithms. Finally, following a 22-month deployment of DLB at Google, where it facilitates large-scale Generative AI inference for thousands of different machine learning models and millions of requests per second, we detail the design choices and practical experiences gained from the system in production. Analysis of production migrations demonstrates that DLB yields statistically significant latency reductions compared to the legacy baseline, including a 17\% decrease in median latency and a 13\% decrease at the p95 tail.

发表机构

  • Google Research(谷歌研究院)
  • Columbia University(哥伦比亚大学)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑