arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32217cs.DC

面向云原生HPC的自适应运行时系统中的无重启弹性机制

No-Restart Elasticity in an Adaptive Runtime System for Cloud-Native HPC

Aditya Bhosale, Laxmikant Kale

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Charm++运行时系统中的无重启重新缩放机制,通过成员协商、通信端点对账和保留状态的初始化路径,将弹性开销从数秒降至毫秒级,并消除GPU检查点依赖,显著提升竞价实例下的HPC弹性效率。

中文摘要 AI 辅助

利用折扣竞价实例进行高性能计算(HPC)要求应用程序在运行时改变其资源分配,在中断前收缩,并在替代容量上扩展。现有的弹性机制将重新缩放实现为完整的进程拆除,随后在新处理器数量下进行冷重启,而重启阶段占总开销的比例高达95%,且随着节点数量增加而增长,在GPU上由于CUDA上下文初始化而大幅增加。在本文中,我们为Charm++运行时系统提出了一种无重启的重新缩放机制,其中幸存的进程永不退出:一次重新缩放操作包括与轻量级外部协调器的成员协商、UCX通信端点与新集群视图的对账,以及返回到保留活动应用程序状态的运行时初始化路径的顶部。这将重新缩放操作的成本(不包括任何重新缩放模型都需要的负载均衡步骤)从数秒降低到CPU上的8-15毫秒和GPU上的7-11毫秒,在4到32个实例范围内。由于进程幸存,GPU设备状态原地保留,消除了以前GPU弹性所需的检查点守护进程,并且与启动器无关的引导机制消除了对受监督进程管理器的依赖,后者与竞价实例中断不兼容。与现有的竞价实例管理框架集成后,该机制将八次同时中断的端到端开销降低到CPU运行时的0.2%和GPU运行时的0.6%,并将作业可在低于1%开销下重新缩放的速率提高了七倍。

英文摘要

Exploiting discounted spot instances for HPC requires an application to change its resource allocation at runtime, shrinking ahead of an interruption and expanding onto replacement capacity. Existing elasticity mechanisms implement rescaling as a full process teardown followed by a cold restart at the new processor count, and the restart stage accounts for up to 95% of the total overhead, growing with node count and increasing substantially on GPUs due to CUDA context initialization. In this paper, we present a no-restart rescaling mechanism for the Charm++ runtime system in which surviving processes never exit: a rescaling operation consists of a membership negotiation with a lightweight external coordinator, reconciliation of the UCX communication endpoints with the new cluster view, and a return to the top of the runtime initialization path that preserves live application state. This reduces the cost of a rescaling operation, excluding the load balancing step that any rescaling model requires, from multiple seconds to 8--15ms on CPUs and 7--11ms on GPUs, at 4 to 32 instances. Because processes survive, GPU device state persists in place, eliminating the checkpointing daemons previously required for GPU elasticity, and a launcher-independent bootstrap mechanism removes the dependence on supervised process managers, which are incompatible with spot instance interruptions. Integrated with an existing spot instance management framework, the mechanism cuts the end-to-end overhead of eight simultaneous interruptions to 0.2% of runtime on CPUs and 0.6% on GPUs, and raises the rate at which a job can be rescaled below 1% overhead by a factor of seven.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑