发表机构
Nanyang Technological University; Cornell University; University of Bristol(南洋理工大学; 康奈尔大学; 布里斯托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PowerZooJax是一个基于JAX的电力系统强化学习基准套件,提供五个受约束的马尔可夫决策过程任务,通过GPU加速实现大规模评估,并标准化评估策略回报、安全违规和分布外压力条件。
AI 中文摘要
电力系统运行是一个安全关键的序贯决策问题,使其成为强化学习(RL)的自然测试平台。然而,现有的电力系统强化学习环境往往范围狭窄,且受限于基于CPU的仿真工作流的计算能力,使得大规模评估变得困难。我们推出了PowerZooJax,一个基于JAX的电力系统运行强化学习基准套件。它提供了五个受约束的马尔可夫决策过程任务,涵盖发电、输电、配电、分布式能源资源和数据中心微电网。通过将潮流计算、经济调度、市场出清和设备动态重写为JAX计算图,PowerZooJax使整个训练和评估循环保持在GPU上。实验表明,与基于CPU的仿真相比,速度显著提升,并展示了策略回报、安全违规和分布外压力条件的标准化评估。我们的开源基准可在以下网址获取:此https URL。
英文摘要
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
Comments42 pages