学习运行电力网络:受AlphaZero启发的有效拓扑控制方法
Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对电网拓扑控制的组合空间与约束问题,研究受AlphaZero启发的模型方法,发现优化的AlphaZero达98.43%生存能力,需结合领域启发式、二元奖励与受限观测空间。
AI中文摘要:
随着波动性可再生能源的整合增加了现代电网的压力,将强化学习(RL)用于自主拓扑重构已成为保持紧张电网稳定和运行的有前景的研究领域。与传统的再调度措施相比,拓扑操作提供了一种更廉价、更具成本效益的电网拥堵管理方式。然而,其实施受到巨大的组合操作空间和严格的运行约束的阻碍。本文研究了基于模型的、受AlphaZero启发的方法的有效性,这些方法利用蒙特卡洛树搜索(MCTS)进行主动电网管理。我们系统评估了奖励函数、观测密度和搜索引导如何影响智能体的生存能力。结果表明,优化后的AlphaZero方法达到了98.43%的峰值生存能力,显著优于近端策略优化(PPO)变体。我们发现,在没有预先学习的策略或价值函数引导的情况下进行MCTS可提高训练效率,且简单的二元生存奖励比复杂的多目标函数提供更有效的搜索引导。我们的研究结果表明,尽管AlphaZero是拓扑控制的强大框架,但纯强化学习并不足够;相反,有效且可靠的系统需要“极简”整合特定领域的启发式规则、二元奖励以及线路负荷的受限观测空间。
英文摘要:
As the integration of volatile renewable energy sources increases the strain on modern power grids, the use of Reinforcement Learning (RL) for autonomous topological reconfiguration has emerged as a promising research field to keep strained grids stable and operational. Compared to traditional redispatching measures, topological actions offer a cheaper and more cost-effective way to manage grid congestion. However, their implementation is hindered by a vast combinatorial action space and strict operational constraints. This paper investigates the effectiveness of model-based AlphaZero-inspired approaches that utilize Monte Carlo Tree Search (MCTS) for proactive grid management. We systematically evaluate how reward functions, observation density, and search guidance influence an agent's survivability. Our results demonstrate that the optimized AlphaZero approach achieves a peak survivability of 98.43%, significantly outperforming the proximal policy optimization (PPO) variant. We find that conducting the MCTS without guidance from a prior learned policy or value function can enhance training efficiency, and that a straightforward binary survival reward provides more effective search guidance than complex, multi-objective functions. Our findings demonstrate that while AlphaZero is a powerful framework for topological control, pure reinforcement learning is not sufficient; rather, an effective and reliable system requires a 'minimalist' integration of domain-specific heuristics, binary rewards, and a restricted observation space of line loads.