一种用于带人工智能数据中心的配电网的预测-调度框架
A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers
浏览论文内容
中文总结 AI 辅助
针对人工智能数据中心电力问题,提出预测-调度(PTS)框架,集成工作负载预测与调度优化,利用可微凸优化及引入工作负载过度转移损失评估调度质量,相比传统基线显著降低运营成本、提高系统安全性。
中文摘要 AI 辅助
人工智能数据中心中GPU密集型工作负载的激增带来了巨大的能源需求,导致成本飙升和当地配电网承受巨大压力。通过精确的工作负载预测来协调容错工作负载调度与电网状况可以缓解这些问题。然而,传统方法存在关键差距,即最小化预测误差不一定能使下游运营损失最小化。因此,本文提出了一种端到端的预测-调度(PTS)框架,将上游工作负载预测与下游调度优化相结合。通过利用可微凸优化,PTS框架将输入特征直接映射到最优调度并实现基于梯度的训练。此外,为了考虑数据中心的容量,引入了一种将电力成本与负载削减惩罚相结合的工作负载过度转移损失来评估调度质量。实验表明,与传统的两阶段基线相比,所提出的框架显著降低了运营成本并提高了系统安全性。
英文摘要
The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.