AI 中文总结
MARS是无需训练的HPC调度器,通过奖励函数配置优化目标,在生产工作负载评估中,其两种变体分别显著降低尾部等待时间、恢复维护前利用率,优于DRL和传统启发式算法。
AI 中文摘要
现代高性能计算(HPC)系统依赖静态启发式算法和人工管理进行作业调度与预留管理。深度强化学习(DRL)虽展现出良好的调度性能,但需要历史训练数据,且在训练时就固定了优化目标,当优先级发生变化时,操作人员必须重新训练模型。本文介绍MARS(Monte Carlo Tree Search-based Adaptive and Responsive Scheduler,基于蒙特卡洛树搜索的自适应响应式调度器),这是一种无需训练的HPC调度器,其优化目标可通过奖励函数配置,而非嵌入到学习模型中。MARS使用轻量级离散事件模拟器,在严格的时间预算内探索调度决策的未来后果,在每个调度周期内适配配置的奖励函数。我们在阿贡领导力计算设施的两个系统的长达一年的生产工作负载上评估MARS:4360节点的Theta和560节点的Polaris,使用两种奖励函数:等待时间最小化(MARS-CW)和利用率最大化(MARS-CU)。与仅对当前队列做出反应或等待回填寻找空隙的DRL和启发式算法不同,MARS利用前瞻性主动排空系统并围绕未来预留进行规划,对系统进行打包以避免预留窗口前常见的碎片化和利用率下降。MARS-CW在Theta上将尾部等待时间减少了64%,在Polaris上较生产WFP启发式算法减少了43%;而MARS-CU则在维护前的48小时内恢复了利用率,证明MARS可通过奖励重配置来实现任一目标。
英文摘要
Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management. Deep Reinforcement Learning (DRL) has shown promising scheduling performance but requires historical training data and fixes the optimization goal at training time, forcing operators to retrain whenever priorities shift. We introduce MARS (Monte Carlo Tree Search-based Adaptive and Responsive Scheduler), a training-free HPC scheduler whose optimization goal is configurable through a reward function rather than baked into a learned model. MARS uses a lightweight discrete-event simulator to explore the future consequences of scheduling decisions within a strict time budget, adapting to the configured reward at each scheduling cycle. We evaluate MARS on year-long production workloads from two systems at Argonne Leadership Computing Facility -- 4,360-node Theta and 560-node Polaris---under two reward functions: wait-time minimization (MARS-CW) and utilization maximization (MARS-CU). Unlike DRL and heuristics, which only react to the current queue or wait for backfill to find holes, MARS exploits look-ahead to proactively drain the system and plan around future reservations, packing the system to avoid the fragmentation and utilization drop that typically precede reservation windows. MARS-CW reduces tail wait time by 64% on Theta and 43% on Polaris over the production WFP heuristic, while MARS-CU recovers utilization in the 48 hours leading into maintenance, demonstrating that MARS can target either objective via reward reconfiguration.
Comments10 pages, 6 figures. Accepted at the 28th IEEE International Conference on Cluster Computing (CLUSTER 2026), September 22-25, 2026, Alexandria, VA, USA. To appear in IEEE Xplore