arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COMPASS-ABS:减少共享GPU集群中深度学习训练工作负载的碎片化

COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

Yukai Zhou, Hongfan Wu

arXiv 2609.18519首次发表:更新:

AI 中文总结

针对共享GPU集群资源碎片化问题,提出基于部分节点度量的SIF指标和COMPASS-ABS调度算法,通过锚点空间限制集群状态,在理论与生产条件下保证SIF有界,实验验证能提高利用率并缩短作业完成时间。

AI 中文摘要

随着深度学习技术的快速发展,共享GPU集群接收到越来越多的深度学习训练(DLT)作业。然而,资源碎片化使得此类集群利用率低下,并迫使在其上运行的DLT作业承受较长的周转时间。已有大量研究致力于量化碎片化并开发缓解其影响的调度算法。然而,现有的碎片化度量在缺乏工作负载分布信息时会失效,而当前的调度器无法持续将资源碎片化维持在较低水平。为解决这些问题,我们首先引入了调度器诱导碎片化(SIF),这是一种基于部分节点概念的度量,不依赖于历史工作负载知识。随后,我们提出了COMPASS-ABS,它采用COMPact-ASSured(COMPASS)算法将集群状态限制在一个紧凑的基于锚点的空间(ABS)内,该空间的构建充分利用了主导工作负载大小与节点容量之间的拓扑对齐。此外,我们还证明了在满足理论与生产相匹配的工作负载组成条件下,SIF被限制在$\ rac{2}{N}$以内。在物理集群和模拟集群上进行的评估表明,COMPASS-ABS在提高资源利用率、通过减少碎片化来缩短DLT作业完成时间方面具有有效性。

英文摘要

With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. Yet resource fragmentation make such clusters underutilized and forces the DLT jobs running on them to endure long turnaround times. Extensive research has been devoted to quantifying fragmentation and developing scheduling algorithms that alleviate its impact. However, existing fragmentation measures break down in the absence of workload distribution information, while current schedulers cannot continuously maintain resource fragmentation at a low level. To tackle these problems, we first introduce Scheduler-Induced Fragmentation (SIF), a metric built on the notion of partial-nodes that is independent of historical workload knowledge. We then propose COMPASS-ABS, which employs the COMPact-ASSured (COMPASS) algorithm to confine the cluster state within a tight Anchor-Based Space (ABS), whose construction fully leverages the topological alignment between dominant workload size and node capacity. Moreover. We also prove that it ensures SIF is bounded by $\frac{2}{N}$ under a workload composition condition that matches both theory and production. Evaluations implemented on a physical cluster and a simulated cluster demonstrate COMPASS-ABS effectiveness at improving resource utilization, reducing DLT job completion time by reducing fragmentation.

Comments22 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑