AI 中文总结
本文提出SNF-ICON调度框架,结合SNF组调度、预测唤醒定时与自适应热备用控制,经多场景实验验证其可降低平均等待时间,且不存在适用于所有场景的最优策略。
AI 中文摘要
高性能计算(HPC)集群的电源状态管理需在减少空闲能耗的同时,避免刚性并行作业出现过度的唤醒延迟。本文提出SNF-ICON,一种事件驱动型控制器,结合最小需求优先(SNF)组调度、预测唤醒定时与自适应热备用控制。在每次调度器调用时,会筛选近期的到达间隔与完成服务样本,判断其是否满足充足性、类指数变异性、低一阶自相关性及可接受的柯尔莫哥洛夫-斯米尔诺夫距离等条件;被拒绝或数据稀疏的窗口使用SNF+IPM(智能电源管理器),被接受的窗口则激活释放预测与指数型下一个事件模型。仅当队列、事件及到达时效性条件允许时,才应用热备用优化,在带超时上限的时间范围内权衡估计等待时间与非计算能耗。我们在AOBA衍生的64节点模型上评估4个DAS2轨迹段与生成的马尔可夫 workload,在AOBA衍生的1152节点模型上评估SDSC Blue。将SNF-ICON与SNF+IPM、带IPM的先来先服务(FCFS)回填算法对比,结果显示其在全部6种场景下均降低了相对于FCFS基线的平均等待时间,且在5种场景下至少接近一种启发式能耗基线;生成的 workload 大量时间处于ICON模式,而DAS2 workload 主要处于 fallback 模式。此外,跨平台结果显示其强烈依赖节点转换与电源模型,因此不存在在所有场景下均最优的单一策略或参数集。
英文摘要
Power-state management in high-performance computing (HPC) clusters must reduce idle energy without excessive wake-up delays for rigid parallel jobs. This paper presents SNF-ICON, an event-driven controller combining smallest-need-first (SNF) gang scheduling, predictive wake timing, and adaptive warm-spare control. At each scheduler invocation, recent interarrival and completed-service samples are screened for sufficiency, exponential-like variability, low lag-one autocorrelation, and acceptable Kolmogorov-Smirnov distance. Rejected or data-sparse windows use SNF+IPM (Intelligent Power Manager), whereas accepted windows activate release prediction and an exponential next-event model. Warm-spare optimization is applied only when queue, event, and arrival-recency conditions permit, balancing estimated waiting and non-compute energy over a timeout-capped horizon. We evaluate four DAS2 trace segments and a generated Markovian workload on AOBA-derived 64-node models, plus SDSC Blue on an AOBA-derived 1152-node model. SNF-ICON is compared with SNF+IPM and First Come First Served (FCFS) + backfilling with IPM. It reduces average waiting time relative to the FCFS-based baseline in all six cases and remains close to at least one heuristic energy baseline in five. The generated workload spends substantial time in ICON mode, whereas DAS2 workloads operate mainly in fallback. Furthermore, cross-platform results show strong dependence on node-transition and power models. Thus, no single policy or parameter set works best in every case.
Comments18 pages, 10 figures