arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07230cs.NIcs.DCcs.PF

Ofan:面向AI训练的最优负载均衡

Ofan: Optimal Load Balancing for AI Training

  • UC Berkeley(加州大学伯克利分校)
  • Technion(以色列理工学院)
  • NVIDIA(英伟达)
  • ICSI(国际计算机科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Sarah McClure, Evyatar Cohen, Jakob Krebs, Alexander Shpiner, Mark Silberstein, Scott Shenker, Sylvia Ratnasamy, Isaac Keslassy

AI总结:

针对AI训练中现有数据包喷洒算法在胖树网络上负载不均导致集体完成时间膨胀的问题,提出目的地感知的交换机负载均衡方案Ofan,实现O(1)排队延迟,在Llama-3 405B模型上减少CCT膨胀16-39倍。

AI中文摘要:

AI工作负载对极端集体完成时间(CCT)的需求挑战了现有的数据包喷洒算法,这些算法在满线速发送工作负载时可能难以高效地进行负载均衡。我们将此归因于一个结构性原因:在胖树网络中,一旦数据包选择了上行路径,其到目的地的下行路径是唯一的,因此对目的地无感知的方案无法消除其造成的不平衡。我们证明此类方案在消息大小为$m$时可能遭受$\Theta(\sqrt{m})$的排队延迟,从而最终触发拥塞控制的速率降低。相反,我们提出了Ofan,一种基于交换机的目的地感知负载均衡方案,可实现$O(1)$的排队延迟。我们还介绍了其pOfan变体,该变体适配当前交换机的流水线架构。我们的P4实现表明其资源消耗适中。使用Llama-3 405B参数模型的端到端FSDP2评估显示,与现有算法相比,Ofan将CCT膨胀减少了$16$--$39$倍。

英文摘要:

The extreme collective completion time (CCT) demands of AI workloads challenge existing packet spraying algorithms, which can have trouble efficiently load-balancing workloads that are sent at full line rates. We trace this to a structural cause: on a fat tree, once a packet picks its upward path, the downward path to its destination is unique, so destination-oblivious schemes cannot undo the imbalance it creates. We prove that such schemes can suffer from $Θ(\sqrt{m})$ queueing for messages of size $m$, thus eventually triggering rate reductions by the congestion control. Instead, we suggest Ofan, a switch-based destination-aware LB scheme that can reach $O(1)$ queueing. We also present its pOfan variant that fits the pipe-based architecture of current switches. Our P4 implementation shows that it consumes modest resources. An end-to-end FSDP2 evaluation with Llama-3 405B-parameter models shows that Ofan cuts CCT inflation by $16$--$39\times$ when compared to existing algorithms.

补充信息

↑