arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ShardMeter:无需猜测的分片与地理分布式训练

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

Tim Beringer, Patrick Diem, Felix Wolf, Arya Mazaheri

arXiv 2608.23840首次发表:更新:

发表机构

Technical University of Darmstadt; PanocularAI(达姆施塔特工业大学; 潘奥库拉AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ShardMeter是一种轻量级分析性能模型,可预测Transformer工作负载在分片分布式训练中的端到端运行时间,助力快速选择近最优部署方案、规避试错成本。

AI 中文摘要

训练大规模AI模型通常超出单个数据中心的承载能力,需要采用分片、多集群及去中心化训练方式。然而,资源分配的巨大空间使得穷尽式基准测试和手动调优难以实现,而性能取决于模型规模、GPU内存、批量大小、带宽和分片策略等紧密耦合的因素。我们提出ShardMeter,这是一种轻量级分析性能模型,可预测基于Transformer的工作负载在任意分片、分布式甚至去中心化训练场景下的端到端运行时间。给定模型特性和目标硬件拓扑,ShardMeter可估算每GPU及每岛吞吐量、训练成本、总挂钟时间,并识别性能瓶颈。我们的分析揭示了岛规模增加时的收益递减机制,量化了计算与通信受限扩展之间的转变,评估了超参数权衡,并为大规模去中心化训练建模了成本-吞吐量关系。ShardMeter将这些见解可视化,可快速探索配置空间,选择近最优部署方案,避免代价高昂的试错。

英文摘要

Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training. However, the huge space of resource allocations makes exhaustive benchmarking and manual tuning impractical, while performance depends on tightly coupled factors like model size, GPU memory, batch size, bandwidth, and sharding strategy. We introduce ShardMeter, a lightweight analytical performance model that predicts the end-to-end runtime of transformer-based workloads across arbitrary sharded, distributed, and even decentralized training. Given a model's characteristics and a target hardware topology, ShardMeter estimates per-GPU and per-island throughput, training cost, total wall-clock time, and identifies performance bottlenecks. Our analysis reveals diminishing-return regimes as island size increases, quantifies transitions between compute- and communication-bound scaling, evaluates hyperparameter trade-offs, and models cost-throughput for large-scale decentralized training. ShardMeter exposes these insights to quickly explore the configuration space, choose near-optimal deployment plans, and avoid costly trial and error.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑