发表机构
Beihang University; University of Leeds(北京航空航天大学; 利兹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多租户GPU集群利用率低和排队延迟高的问题,提出DeepShare调度器,通过连续租户保障信号协调配额借用、抢占与共享,实现利用率提升29.5%、延迟降低46%。
AI 中文摘要
多租户GPU集群即使在租户经历长时间排队延迟时也常常利用不足,这是因为配额控制、队列排序、抢占和GPU共享由不同的本地信号驱动。我们提出DeepShare,一种调度器,它使用连续的租户保障信号在运行时协调这些决策。DeepShare结合了弹性配额借用、租户特定运行时预测、成本感知的尽力而为抢占和干扰感知的MPS共存,同时使用相同的保障信号来决定何时应回收借用容量以及何时共享应变得更加保守。在基于轨迹的实验中,涉及23,859个Venus作业和3,200个内部作业,DeepShare实现了70.58%的平均GPU利用率,比最强的非侵入式共享基线提高了29.5%,同时将平均排队延迟降低了46%。在16-GPU Kubernetes测试平台上,它将平均作业完成时间减少了34%,并为有保障租户维持了93%的QoS合规性。这些结果表明,将租户保障视为运行时控制回路,比独立优化配额、调度和资源共享能实现更有利的利用率-QoS权衡。
英文摘要
Multi-tenant GPU clusters frequently remain underutilized even when tenants experience long queueing delays, because quota control, queue ordering, preemption, and GPU sharing are driven by different local signals. We present DeepShare, a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime. DeepShare combines elastic quota borrowing, tenant-specific runtime prediction, cost-aware best-effort preemption, and interference-aware MPS colocation, while using the same assurance signal to decide when borrowed capacity should be reclaimed and when sharing should become more conservative. In trace-driven experiments on 23,859 Venus jobs and 3,200 internal jobs, DeepShare achieves an average GPU utilization of 70.58%, a 29.5% improvement over the strongest non-intrusive sharing baseline, while reducing average queueing delay by 46%. On a 16-GPU Kubernetes testbed, it reduces the average job completion time by 34% and maintains 93% QoS compliance for guaranteed tenants. These results show that treating tenant assurance as a runtime control loop achieves a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Comments13 pages. Accepted at IEEE CLUSTER 2026