发表机构
UNIST; POSTECH; UC Merced; USC(蔚山科学技术院; 浦项工科大学; 加州大学默塞德分校; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Arachne提出基于学习的规划器,将并行训练规划简化为流水线结构搜索,在动态异构GPU集群上快速生成高吞吐量方案,吞吐量最高提升84.5%或4.6倍。
AI 中文摘要
在共享GPU基础设施上训练大型机器学习模型面临两个挑战:(1)GPU可用性会随着租户不同的资源需求而动态变化,(2)随着数据中心不断采用新的GPU代际,硬件异构性不断累积。由于异构GPU类型和节点规模带来的巨大搜索空间,训练规划器必须激进地剪枝以保持可处理性,同时必须在集群配置变化时迅速生成高吞吐量的规划。Arachne通过一种基于学习的规划器实现了这一目标,该规划器将完整的规划问题简化为对流水线结构的搜索。Arachne将规划决策封装在结构模板中,并离线学习如何从模板中为多样化的集群配置构建规划。这种设计之所以有效,是因为结构决策构成了并行规划中性能关键的核心,而一旦规划结构固定,其余部分可以通过规则或从小的定价候选集中得出。评估表明,在具有不同GPU类型和节点规模的集群上,针对三种不同规模的模型,Arachne匹配或超过了五种现有规划器找到的最佳规划,在稠密模型上吞吐量最高提升84.5%,在MoE模型上最高提升4.6倍。
英文摘要
Training large machine learning models on shared GPU infrastructures faces two challenges: (1) GPU availability shifts dynamically with varying resource demands from tenants, and (2) hardware heterogeneity accumulates as datacenters continuously adopt new GPU generations. Due to the vast search space induced by heterogeneous GPU types and node sizes, training planners must prune it aggressively to remain tractable, yet must also derive high-throughput plans promptly as cluster configurations change. Heddle achieves this goal through a learning-based planner that reduces the full planning problem to a search over pipeline structures. Heddle encapsulates planning decisions in a structural template and learns to construct plans from templates over diverse cluster configurations offline. This design is effective because structural decisions constitute the performancecritical core of a parallelism plan, while the rest follows by rule or from a small priced candidate set once the plan structure is fixed. Evaluation shows that Heddle matches or exceeds the best plan found by five existing planners across clusters with varying GPU types and node sizes for three models of different sizes by up to 84.5% in throughput on dense models and 4.6x on MoE models.