arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨设施 HPC 上的 LLM 预训练:弹性聚合、数据租用与队列感知放置

Cross-Facility LLM Pre-training on HPC: Elastic Aggregation, Data Leasing, and Queue-Aware Placement

Zarè Palanciyan, Thomas van Osch, Douwe van der Wal, Olivera Kotevska, Tim Kok

arXiv 2610.03457首次发表:更新:

发表机构

SURF; Oak Ridge National Laboratory(SURF; 橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一个跨设施HPC预训练系统,通过弹性聚合、数据租用和队列感知放置,在三台相距7400公里的超级计算机上训练语言模型,以可测量成本实现碎片化分配的统一训练。

AI 中文摘要

学术计算是碎片化的:分配按设施授予,设施在加速器和软件栈上各不相同,独立调度作业,并且既不共享网络也不共享文件系统。我们提出了一个系统,汇集此类分配,以在跨越两大洲、相距最远 7400 公里的三台超级计算机上预训练单个语言模型:Snellius(NVIDIA H100)、LUMI 和 Frontier(均为 AMD MI250X)。该系统结合了 (i) DiLoCo 风格的双循环训练,采用弹性、令牌加权的 Nesterov 外部步骤,其中零个、一个或多个活跃站点均为正常状态;(ii) DARL,一种数据租用协议,其基于心跳的租约保证在一个周期内,任何样本在崩溃、延迟加入和工作窃取情况下都不会被训练两次或丢失;(iii) 队列感知放置,由无特权的 sbatch --test-only 探测驱动。在 C4 上训练 Qwen3-0.6B 模型 20,000 个优化器步骤,三站点运行达到的保留困惑度为 34.7,而集中式基线为 28.2。每轮开销(权重交换和检查点)保持在接近 110 秒,无论本地步骤数 H 为多少,因此其在墙钟时间中的占比从 H=100 时的 32% 降至 H=1,000 时的 5.7% 和 H=2,000 时的 3.1%。在一次 23.8 小时的三站点运行中,有四个站点离开,1.2% 的已授予数据块被回收,且没有重复或丢失。在理想化的队列模型投影中,与等待所有站点分配相比,队列感知放置将达成目标的时间缩短了 18-43%。因此,跨站点预训练是可操作的而非竞争性的:它以可测量的成本将碎片化分配转化为一次训练运行。

英文摘要

Academic compute is fragmented: allocations are granted per facility, and facilities differ in accelerators and software stacks, schedule jobs independently, and share neither a network nor a filesystem. We present a system that pools such allocations to pre-train a single language model across three supercomputers on two continents, up to 7,400km apart: Snellius (NVIDIA H100), LUMI and Frontier (both AMD MI250X). It combines (i) DiLoCo-style two-loop training with an elastic, token-weighted Nesterov outer step for which zero, one or many live sites are all normal states; (ii) DARL, a data-leasing protocol whose heartbeat-backed leases guarantee that, within an epoch, no sample is trained twice or lost under crashes, late joins and work stealing; and (iii) queue-aware placement driven by unprivileged sbatch --test-only probes. Training Qwen3-0.6B on C4 for 20,000 optimizer steps, the three-site run reaches a held-out perplexity of 34.7, against 28.2 for a centralised baseline. Per-round overhead (weight exchange and checkpointing) stays near 110 s regardless of the number of local steps H, so its share of wall-clock time falls from 32% at H=100 to 5.7% at H=1,000 and 3.1% at H=2,000. In a 23.8 h three-site run with four site departures, 1.2% of granted data blocks were reclaimed and none was duplicated or lost. In an idealized queue-model projection, queue-aware placement shortens time-to-target by 18-43% compared with waiting for all sites to be allocated. Cross-site pre-training is thus operational rather than competitive: it turns fragmented allocations into one training run at a measured cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑