arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26453cs.LGcs.DCcs.NI

基于智能网络的分布式训练

Distributed Training using an Intelligent Network

Nihar Shah, Ben Blier

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对广域网分布式训练的带宽、延迟等瓶颈,提出结合多播、在线FPGA的系统方案与基于拓扑的旋转团同步调度算法,在九城市DoubleZero网络拓扑上验证可缩小与同位置训练的性能差距。

中文摘要 AI 辅助

跨广域网(WAN)的分布式训练颇具挑战性,因为计算岛之间的持续参数交换会受限于有限的带宽、高延迟以及不均匀的拓扑结构。我们提出让网络成为训练的主动参与者。在系统层面,这类网络应利用(i)多播技术复制出站流量,以及(ii)在线FPGA聚合入站流量,以缓解出口和入口瓶颈。这些技术已用于数据中心内工作节点间的训练,本文将其扩展至广域网场景。在算法层面,我们开发了一个优化框架,可围绕底层网络拓扑和上述技术生成丰富的同步调度方案(即计算岛的旋转团),以最大化信息交换。最后,我们在基于DoubleZero网络(一个配备上述两种技术的实时可编程广域网)构建的九城市拓扑上验证了该方案,展示了最优调度如何随网络能力变化。这些技术结合可缩小与同位置训练黄金标准的差距。

英文摘要

Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwidth, high latency, and uneven topology. We propose making the network an active participant in training. On the systems side, such networks should leverage (i) multicast technology to replicate outbound traffic and (ii) in-line FPGAs to aggregate inbound traffic, to ease egress and ingress bottlenecks. These technologies are used for training across workers within a data center, but this paper extends them to the WAN. On the algorithms side, we develop an optimization framework that produces rich synchronization schedules (namely, rotating cliques of islands) around the underlying network topology and these technologies, to maximize information exchange. Finally, we illustrate this on a nine-city topology modeled on the DoubleZero network, a live programmable WAN equipped with both technologies, and show how the optimal schedules shift with the network's capabilities. Together, these can narrow the gap to the gold standard of colocated training.

↑