arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OrigaMIG:基于邻域受限BILP与实时迁移的MIG感知虚拟机放置

OrigaMIG: MIG-Aware VM Placement with a Neighborhood-Restricted BILP and Live Migration

Ahmad Siavashi, Mahmoud Momtazpour

arXiv 2610.06646首次发表:更新:

发表机构

Amirkabir University of Technology(阿米尔卡比尔理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MIG感知的虚拟机放置问题,提出OrigaMIG优化器,采用贪心与邻域受限BILP结合实时迁移,最大化请求接受率、整合资源并降低迁移开销,实验证明其优于现有策略。

AI 中文摘要

云计算中GPU的广泛使用,因大型语言模型(LLM)服务的普及而加速,加之多租户需求的日益增长,推动了高效GPU资源管理创新解决方案的发展。NVIDIA的多实例GPU(MIG)技术通过提供隔离实例(以MIG支持的虚拟GPU(vGPU)形式提供)来实现云数据中心中的共享GPU使用。然而,MIG放置规则常常导致碎片化和次优的资源分配。在本工作中,我们将MIG感知的虚拟机(VM)放置问题正式建模为一个二进制整数线性规划(BILP)问题,旨在最大化请求接受率、整合资源并减少迁移开销。基于此公式,我们提出了OrigaMIG,一个MIG感知的放置优化器。OrigaMIG使用轻量级贪心规则立即放置每个到达的请求,仅当请求无法放置或某类GPU变得碎片化时才调用求解器。每次调用在物理机(PM)的小邻域上求解BILP,保持数据中心其余部分固定,并将求解结果作为实时迁移执行。我们在模拟中将OrigaMIG与默认放置策略及两种最先进的MIG感知策略进行比较,涵盖六种负载、十一种工作负载和硬件变体,以及最多4096台PM的数据中心。在128台PM上的256个A100和A30 GPU中,OrigaMIG在30次配对运行中的29次中比更强的策略保持更少的活跃PM,同时迁移的VM数量减少29%至49%,并且在每个负载下每接纳GPU小时的能耗最低。与默认放置相比,它保持的活跃PM最多减少12.5%,接纳的请求GPU内存最多增加7.8个百分点。在小型数据中心中,它与最优活跃PM数量的差距在4.1%以内。

英文摘要

The extensive use of GPUs in cloud computing, accelerated by the spread of large language model (LLM) services, and the growing need for multitenancy have driven the development of innovative solutions for efficient GPU resource management. Multi-Instance GPU (MIG) technology from NVIDIA enables shared GPU usage in cloud data centers by providing isolated instances, which are offered as MIG-backed virtual GPUs (vGPUs). However, MIG placement rules often lead to fragmentation and suboptimal resource allocation. In this work, we formally model the MIG-aware virtual machine (VM) placement as a binary integer linear programming (BILP) problem aimed at maximizing request acceptance, consolidating resources, and reducing migration overhead. Building upon this formulation, we propose OrigaMIG, a MIG-aware placement optimizer. OrigaMIG places each arriving request at once with a lightweight greedy rule and calls the solver only when a request cannot be placed or the GPUs of a class become fragmented. Each call solves the BILP on a small neighborhood of physical machines (PMs), keeps the rest of the data center fixed, and executes the solution as live migrations. We compare OrigaMIG with the default placement and with two state-of-the-art MIG-aware policies in simulation, across six loads, eleven variants of the workload and the hardware, and data centers of up to 4096 PMs. On 256 A100 and A30 GPUs in 128 PMs, OrigaMIG keeps fewer PMs active than the stronger policy in 29 of 30 paired runs while migrating 29% to 49% fewer VMs, and it uses the least energy per admitted GPU-hour at every load. Against the default placement, it keeps up to 12.5% fewer PMs active and admits up to 7.8 percentage points more of the requested GPU memory. On small data centers, it stays within 4.1% of the optimal number of active PMs.

Comments19 pages, 8 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑