arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于人工智能数据中心的分层半马尔可夫负载模型:将作业调度与批量同步并行功率动态相结合

A Hierarchical Semi-Markov Load Model for AI Data Centers Coupling Job Scheduling with Bulk-Synchronous-Parallel Power Dynamics

Chandan Chaudhary, Atri Bera, Cody Newlun, Mohammed Ben-Idris, Joydeep Mitra

arXiv 2607.12222首次发表:更新:

发表机构

Michigan State University; Sandia National Laboratories(密歇根州立大学; 桑迪亚国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人工智能数据中心负载模型,提出分层半马尔可夫数据中心(HSM-DC)负载模型,该模型在两个时间尺度上耦合两层,能匹配平均功率、分布及峰均比等,强调电网规划需考虑作业到达和调度过程。

AI 中文摘要

人工智能数据中心作为一种主要的新负载类型正在兴起,其功率动态与传统工业负载有根本区别。在训练作业中,批量同步并行算法使每个节点经历计算、同步和检查点步骤,在数秒内使功率在满负荷和接近空闲之间波动。在整个设施中,作业到达、占用节点数小时至数天然后离开,忙碌节点数量每日、每周和每年都在变化。仅关注作业内行为并将设施视为固定忙碌节点集的模型会平滑这些波动并错过真正的峰均比。本文开发了一种分层半马尔可夫数据中心(HSM-DC)负载模型,它在两个时间尺度上耦合两层。作业调度层通过由每日、每周和季节模式塑造的非齐次复合泊松过程创建作业,为每个作业赋予重尾节点数和长度,并按先来先到顺序将作业放置在固定节点池上。作业内层面使每个忙碌节点通过用于BSP步骤的五状态半马尔可夫链,并带有基于状态的奥恩斯坦 - 乌伦贝克噪声。设施功率来自不断变化的节点数和每个节点的功率,设置为与测量的节点数据和设施的直线功率与负载曲线相匹配。在相同规模下配置到参考设施,该模型在负载水平上匹配平均功率、其分布以及峰均比,拟合分数分别为0.9997、0.92和0.82。它在高负载下还能将排队作业的份额匹配到相差一个百分点以内。设施范围内的波动和峰值需求来自作业的到达和调度方式,因此电网规划必须对该过程进行建模,而不仅仅是放大单个节点的功率曲线。

英文摘要

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so the number of busy nodes changes daily, weekly, and yearly. This slower shift drives facility-wide swings and the peak demand that sets the size of the grid link. A model that looks only at within-job behavior, and treats the facility as a fixed set of busy nodes, smooths out these swings and misses the true peak-to-average ratio. This paper develops a hierarchical semi-Markov Data-Center (HSM-DC) load model that couples two layers across two timescales. A job-scheduling layer creates jobs through a non-homogeneous compound-Poisson process shaped by daily, weekly, and seasonal patterns, gives each job a heavy-tailed node count and length, and places jobs on a fixed pool of nodes on a first-come basis. A within-job layer moves each busy node through a five-state semi-Markov chain for the BSP steps, with state-based Ornstein-Uhlenbeck noise. Facility power comes from this changing node count and the per-node power, set to match measured node data and the facility's straight-line power-versus-load curve. Configured to the reference facility at the same scale, the model matches mean power, its spread, and the peak-to-average ratio across load levels, with fit scores of 0.9997, 0.92, and 0.82. It also matches the share of queued jobs to within one point at high load. Facility-wide swings and peak demand come from how jobs arrive and get scheduled, so grid planning must model that process, not just scale up a single node's power curve.

Comments2026 CIGRE Grid of the Future Symposium, October 26-29, 2026 Richmond, Virginia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑