克服百亿亿次规模下的编排瓶颈:一种用于模拟 - 人工智能集成的去中心化、策略驱动方法
Overcoming Orchestration Bottlenecks at Exascale: A Decentralized, Policy-Driven Approach for Sim-AI Ensembles
浏览论文内容
中文总结 AI 辅助
针对科学计算中耦合模拟 - 人工智能工作流的编排瓶颈问题,提出了具有去中心化控制平面和可编程调度策略接口的EnsembleLauncher编排器,在Aurora超级计算机上性能出色,还展示了调度策略对资源利用的显著影响。
中文摘要 AI 辅助
科学计算正日益从单一应用转向由具有不同硬件、规模和运行时要求的高度异构任务组成的耦合模拟 - 人工智能工作流。随着这些工作流扩展到领导级系统,由此产生的极端集成规模和任务变异性会造成编排瓶颈。系统级调度器吞吐量有限,工作流工具因控制平面拓扑僵化和静态调度启发式方法而面临可扩展性问题。我们引入了EnsembleLauncher,一种用于百亿亿次系统的递归分层工作流编排器,具有完全去中心化的控制平面和可编程调度策略接口。在Aurora超级计算机上,EnsembleLauncher成功扩展到整台机器,处理多达八百万个串行任务,性能比现有工具高出四倍多。此外,我们实现了可编程调度接口,并证明调度策略对高变异性集成和代表现代耦合模拟 - 人工智能工作流的主动学习管道的资源利用有重大影响。
英文摘要
Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requirements. As these workflows scale to leadership-class systems, the resulting extreme ensemble sizes and task variability can create orchestration bottlenecks. System-level schedulers are often configured for limited throughput, while workflow tools face scalability issues due to rigid control-plane topologies and static scheduling heuristics. We introduce EnsembleLauncher, a recursively hierarchical workflow orchestrator for exascale systems, featuring a fully decentralized control plane and a programmable scheduling policy interface. On the Aurora supercomputer, EnsembleLauncher successfully scales to the entire machine with up to eight million serial tasks, outperforming state-of-the-art tools by more than four times. Additionally, we implement a programmable scheduling interface and demonstrate a significant impact of scheduling policies on resource utilization for high-variance ensembles and active learning pipelines representative of modern coupled simulation-AI workflows.