arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多租户GPU集群的队列论准入控制

Queue-Theoretic Admission Control for Multi-Tenant GPU Clusters

Sohan Kunkerkar

arXiv 2607.28223首次发表:更新:

AI 中文总结

该研究针对GPU集群准入等待时间不可预测问题,将其建模为多类多资源排队网络,提出基于队列论的准入控制方法,在Kueue系统上验证了方法的有效性。

AI 中文摘要

GPU集群运营商无法预测待处理工作负载的准入等待时长,现有系统采用无正式等待时间保证的贪心启发式算法。我们将GPU集群准入形式化为多类多资源排队网络,证明其结构分解:待处理队列分为可配额工作负载(稳定状态下等待时间有界)与不可行工作负载(未重新配置时无有限边界)。对可配额工作负载,我们将每个集群队列建模为M/G/k系统,其中有效服务器数量k由向量装箱归约确定;在显式随机支配假设下,我们建立O(1/(1-ρ))的等待时间缩放关系。通过向量装箱归约,我们证明多维度资源需求下的最优准入排序是NP难问题。我们在标准Kubernetes工作负载排队系统Kueue上,利用CPU、内存及GPU(通过动态资源分配)资源进行验证,结果显示有效k_eff向量可准确识别瓶颈资源维度,利特尔定律精确成立,爱尔兰C近似值始终以保守方向高估观测到的等待时间。

英文摘要

GPU cluster operators cannot predict how long pending workloads will wait for admission. Existing systems use greedy heuristics with no formal wait time guarantees. We formalize GPU cluster admission as a multi-class, multi-resource queueing network and prove a structural decomposition: the pending queue partitions into quotable workloads (bounded wait time under stability) and unfeasible workloads (no finite bound without reconfiguration). For quotable workloads, we model each cluster queue as an M/G/k system where the effective server count k is determined by a vector packing reduction; under an explicit stochastic domination assumption, we establish O(1/(1-rho)) wait time scaling. We prove that optimal admission ordering is NP-hard under multi-dimensional resource demands via reduction from vector bin packing. We validate on Kueue, the standard Kubernetes workload queuing system, using CPU, memory, and GPU (via Dynamic Resource Allocation) resources. The vector k_eff correctly identifies bottleneck resource dimensions, Little's Law holds exactly, and the Erlang-C approximation consistently overestimates observed wait times in the conservative direction.

Comments9 pages, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑