XMPIaaS:通过协作式进程迁移实现云原生MPI
XMPIaaS: Towards Cloud Native MPI via Cooperative Process Migration
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出XMPIaaS,一种面向云环境的MPI协作式进程迁移系统,通过选择性迁移受抢占影响的进程组,避免全局检查点开销,实现低成本、弹性的云原生MPI执行。
中文摘要 AI 辅助
消息传递接口(MPI)三十年来一直是高性能计算(HPC)的主导编程模型,随着HPC工作负载日益迁移到云基础设施以寻求可扩展性和成本效益,MPI应用必须应对与传统超级计算机根本不同的执行环境:临时资源、动态定价和可抢占实例。在这种不稳定的环境中,无需重启作业即可在节点间迁移运行中的MPI进程的能力,对于实现成本效益高、弹性强的执行至关重要。现有方法要么需要从全局检查点重启整个作业,要么以过高的复杂性透明拦截整个MPI栈。为解决这些挑战,我们提出\ ame,一个面向MPI的协作式迁移系统,支持按需进行选择性进程组迁移。当云实例被调度抢占时,仅受影响的任务(rank)被迁移,而其余进程短暂静止并在原地恢复,从而避免整个作业检查点的开销。\ ame通过MPI进程管理运行时与rank进程之间的协作协议解决此问题。我们基于MPI Sessions API暴露了一个\ exttt{XMPI_quiesce}接口,允许应用标记安全迁移点,并扩展Hydra进程管理器以编排完整的迁移生命周期:rank静止、CRIU检查点/恢复、目标节点上的代理重启以及无缝的rank重连。我们评估并表明,协作式静止阶段占总迁移停机时间的比例不到1.4%,且该停机时间仅由迁移节点的rank数量决定,与作业规模无关,并且该插桩在正常执行期间不引入可测量的开销。
英文摘要
Message Passing Interface (MPI) has been the dominant programming model for High Performance Computing (HPC) for three decades, and as HPC workloads increasingly migrate to cloud infrastructure for scalability and cost efficiency, MPI applications must contend with an execution environment fundamentally unlike traditional supercomputers: ephemeral resources, dynamic pricing and preemptable instances. In such a volatile setting, the ability to relocate running MPI processes between nodes without restarting the job is a necessity for cost-effective, resilient execution. Existing approaches either require restarting the entire job from a global checkpoint, or transparently intercepting the full MPI stack at prohibitive complexity. To address these challenges, we propose \name, a cooperative migration system for MPI that enables selective process group migration on-the-fly. When a cloud instance is scheduled for preemption, only the affected ranks are relocated while the remaining processes briefly quiesce and resume in place, avoiding the cost of a full-job checkpoint. \name tackles this through a cooperative protocol between the MPI process management runtime and rank processes. We expose an \texttt{XMPI\_quiesce} interface built atop the MPI Sessions API that allows applications to mark safe migration points, and we extend the Hydra process manager to orchestrate the full migration lifecycle: rank quiescence, CRIU checkpoint/restore, proxy relaunch on the target node, and seamless rank reconnection. We evaluate and show that the cooperative quiesce phase accounts for less than 1.4\% of total migration downtime, and that this downtime is governed by the migrating node's rank count alone, independent of job size, and the instrumentation introduces no measurable overhead during normal execution.