arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OAK:共享GPU集群上分布式机器学习的重启与年龄感知调度

OAK: Restart- and Age-Aware Scheduling for Distributed Machine Learning on Shared GPU Clusters

Khaled Aljbab, Amine Barrak

arXiv 2609.19024首次发表:更新:

AI 中文总结

OAK通过年龄键和分解重启因子,在共享GPU集群上实现重启与年龄感知调度,显著降低故障下的作业完成时间,且决策延迟极低。

AI 中文摘要

分布式机器学习越来越多地运行在共享的多租户GPU集群上,其中争用和故障是常态。以有效吞吐(goodput)为驱动的调度器最大化瞬时吞吐,但将累计等待时间和重启成本视为次要信号:长时间等待的作业被反复推迟,且每次中断后支付的数分钟检查点加载时间并未被纳入分配决策。我们提出OAK,一个每轮混合整数线性调度器,其复合效用函数在有效吞吐基础上增加了两个一等项:一个年龄键(age key),用于提升具有高累计等待时间的作业;以及一个分解的重启因子,将有效训练时间和检查点开销作为分别测量的量进行追踪,而非单一的聚合比率。我们在一个基于轨迹的模拟器中,于12-GPU集群上针对四个数据并行工作负载,在受控的泊松故障注入下,将OAK与四个代表性的有效吞吐驱动和公平性驱动的基线进行比较,并在一个4-V100真实硬件集群上复现了故障模式下的改进。在无故障条件下,OAK与最强的有效吞吐驱动基线相差在4%以内。在故障率λ=0.1的情况下,相对于其有效吞吐驱动的基础,OAK将平均作业完成时间(JCT)降低了33.9%-58.3%,并将最坏情况JCT降低了58%-76%;仅分解的重启因子一项,相对于先前有效吞吐驱动调度器中使用的聚合重启估计器,就贡献了40%-57%的平均JCT改进和52%-76%的尾部JCT改进。每轮决策延迟为7.65毫秒,比默认设置下的进化搜索替代方案快三个数量级。

英文摘要

Distributed machine learning increasingly runs on shared multi-tenant GPU clusters where contention and failures are routine. Goodput-driven schedulers maximise instantaneous throughput but treat accumulated waiting time and restart cost as second-class signals: long-waiting jobs are repeatedly deferred, and the minutes of checkpoint loading paid after each interruption are not folded back into allocation decisions. We present OAK, a per-round mixed-integer linear scheduler whose composite utility adds two first-class terms to goodput: an age key that elevates jobs with high cumulative waiting time, and a decomposed restart factor that tracks productive training time and checkpoint overhead as separately measured quantities rather than a single aggregate ratio. We evaluate OAK against four representative goodput- and fairness-driven baselines in a trace-driven simulator on a 12-GPU cluster across four data-parallel workloads with controlled Poisson failure injection, and reproduce the failure-mode improvement on a 4-V100 real-hardware cluster. Under failure-free conditions OAK matches the strongest goodput-driven baseline within 4%. Under failures at rate $λ= 0.1$ it reduces mean job completion time (JCT) by 33.9-58.3% and worst-case JCT by 58-76% over its goodput-driven foundation; the decomposed restart factor alone accounts for a 40-57% mean and 52-76% tail JCT improvement against an aggregate restart estimator used in prior goodput-driven schedulers. Per-round decision latency is 7.65 ms, three orders of magnitude faster than evolutionary-search alternatives at default settings.

Journal refThe 32nd IEEE International Conference on Parallel and Distributed Systems (ICPADS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑