AI 中文总结
Poseidon提出基于DAG的显式训练时间模型,结合阶段级剪枝与脊状分布映射,高效搜索异构集群上的LLM预训练并行方案,吞吐量提升达2.76倍。
AI 中文摘要
随着加速器技术的快速发展,在异构加速器集群上预训练大语言模型(LLM)对于最大化硬件利用率变得越来越关键。然而,现有系统存在训练时间建模不准确的问题,这削弱了基于该建模的并行化优化。此外,对于当前方法而言,庞大的配置搜索空间使得穷举搜索不可行,迫使在搜索时间与训练效率之间进行权衡。为克服这些局限,我们提出了Poseidon,一个高效且可扩展的、具有异构感知能力的LLM训练框架。其核心是一个基于有向无环图(DAG)的显式训练时间模型。在此图的基础上,Poseidon采用了两种高效且具有理论依据的策略:通过带部分估计的早停进行阶段级剪枝,以及利用脊状分布模式的层到阶段映射。这些策略在不牺牲最优训练效率的前提下缩减了搜索空间。在异构集群上的实验表明,与最先进的系统相比,Poseidon将训练吞吐量提升了高达2.76倍。
英文摘要
With the rapid advancement of accelerator technologies, pre-training large language models (LLMs) on heterogeneous accelerator clusters has become increasingly crucial for maximizing hardware utilization. Existing systems, however, suffer from inaccurate training time modeling, which undermines the parallelization optimizations built upon it. Moreover, for current approaches, the vast configuration search space makes exhaustive exploration infeasible, forcing a trade-off between search time and training efficiency. To overcome these limitations, we introduce Poseidon, an efficient and scalable LLM training framework designed with heterogeneity awareness. Its core is an explicit training time model based on a directed acyclic graph. Building on this graph, Poseidon employs two efficient, theoretically grounded strategies: stage-level pruning via early stopping with partial estimation, and layer-to-stage mapping exploiting a ridge-like distribution pattern. These strategies reduce the search space without sacrificing optimal training efficiency. Experiments on heterogeneous clusters show that Poseidon improves training throughput by up to $2.76\times$ over state-of-the-art systems.
Comments21 pages, 24 figures