AI 中文总结
研究针对大规模AI训练跨电力受限站点的地理分布式训练能耗高问题,提出分层聚合系统PowerScale,利用广域网延迟层次结构,采用同步-异步同步模式及自适应策略,评估显示其在准确性时间相当或提升时能耗降低达3.9倍。
AI 中文摘要
大规模人工智能训练的功耗需求日益超出单个数据中心的承载能力,使得跨电力受限站点的地理分布式训练成为现实所需。先前工作主要针对准确性时间优化此类训练,采用单层聚合,各站点在每次同步轮次通过广域网直接与中央聚合器交换模型更新,未考虑收敛所需能量。单层聚合本质上能源效率低下。为解决这些低效问题,我们提出PowerScale,一种利用广域网延迟层次结构的分层聚合系统。PowerScale将站点组织成区域集群并应用同步-异步同步模式,基于网络 proximity 和电力可用性形成集群,并使用自适应同步策略。我们在基于Flower的模拟环境中对100个站点规模的PowerScale进行评估,其在准确性时间上与单层基线相当或略有提升,同时能耗降低达3.9倍。
英文摘要
The power demands of large-scale AI training increasingly exceed the capacity of any single data center, making geo-distributed training across power-constrained sites a practical necessity. Prior work optimizes such training mainly for time-to-accuracy using single-tier aggregation, where every site exchanges model updates directly with a central aggregator over the WAN each synchronization round, without accounting for the energy required to reach convergence. Single-tier aggregation is fundamentally energy-inefficient because synchronization barriers force faster sites to idle, full WAN updates dominate communication energy at scale, and fixed synchronization frequency keeps paying the same communication cost even when updates shrink late in training. To address these inefficiencies, we present PowerScale, a hierarchical aggregation system that exploits the latency hierarchy of wide-area networks. PowerScale organizes sites into regional clusters and applies a Sync-Async synchronization modality: sites synchronize frequently with a nearby cluster aggregator over fast local links, while cluster aggregators push pre-aggregated updates asynchronously to a global aggregator over the WAN. PowerScale forms clusters based on both network proximity and power availability, and uses an adaptive synchronization policy that reduces communication energy by adjusting how often clusters synchronize to training progress. This structure shortens synchronization barriers and replaces per-site WAN transmissions with fewer, pre-aggregated transmissions at a lower frequency. We evaluate PowerScale at 100-site scale in a Flower-based simulation environment. PowerScale matches or slightly improves time-to-accuracy compared with single-tier baselines while reducing energy consumption by up to 3.9x.