AI 中文总结
ElastiCo是一种基于Kubernetes的弹性共置框架,通过资源形态转换、弹性影子定价、感知干扰共置三种机制,实现训练与推理工作负载安全共享GPU,显著降低平均JCT、提升集群吞吐量与GPU利用率。
AI 中文摘要
现代GPU集群需同时服务深度学习训练与离线大语言模型推理工作负载,但现有调度器将二者视为孤立的资源消费者,采用僵化的静态分配方式,导致大量GPU容量未被充分利用:训练作业尽管存在周期性空闲阶段仍预留整个设备,而离线推理任务尽管需求模式具有突发性仍过度配置GPU。本文提出ElastiCo,一种弹性共置框架,通过三种集成机制使训练与推理工作负载可安全共享GPU:其一,资源形态转换将每个作业呈现为一系列可行的资源-性能配置;其二,弹性影子定价法通过动态的各资源影子价格,将多资源分配问题分解为各作业的配置选择子问题;其三,感知干扰的共置机制使用基于硬件计数器和任务级特征训练的预测器,估算GPU共享时的成对性能下降。ElastiCo作为原生Kubernetes中间件实现,无需用户代码修改,在64-GPU测试床及大规模轨迹驱动模拟(最高512 GPU)上评估,可将平均JCT降低最高2.94倍,集群吞吐量提升2.02倍,GPU利用率从约25%提升至46%。
英文摘要
Modern GPU clusters must simultaneously serve deep learning training and offline large language model inference workloads, yet existing schedulers treat these as isolated resource consumers with rigid, static allocations. This leaves substantial GPU capacity underutilized: training jobs reserve entire devices despite periodic idle phases, while offline inference tasks over-provision GPUs despite bursty demand patterns. We present ElastiCo, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms. First, Resource Shape Transformation exposes each job as a family of feasible resource-performance configurations. Second, Elastic Shadow Pricing decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices. Third, Interference-Aware Co-location uses a predictor trained on hardware-counter and task-level features to estimate pairwise performance degradation under GPU sharing. Implemented as native Kubernetes middleware requiring no user-code modifications, ElastiCo is evaluated on a 64-GPU testbed and through large-scale trace-driven simulations (up to 512 GPUs), reducing the average JCT by up to 2.94x, increasing the cluster throughput by 2.02x, and increasing the GPU utilization from approximately 25% to 46%.
Comments27 pages, 9 figures, 9 tables. Submitted to Performance Evaluation