发表机构
Friedrich-Alexander-Universität Erlangen–Nürnberg; Ostbayerische Technische Hochschule Amberg-Weiden(埃尔朗根-纽伦堡弗里德里希-亚历山大大学; 东巴伐利亚安贝格-魏登应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用离线强化学习(CQL)基于CIGRE中压网络静态轨迹实现配电网线路选择性跳闸,在225个保留情节中达到F1分数0.9738,并发现密集预测性能与终端保护行为可能导致不同模型排名。
AI 中文摘要
数据驱动的保护方法可以补充传统继电器在配电网中的应用,这些配电网的运行条件会随着分布式发电、开关事件和短路水平的变化而改变。我们利用离线强化学习,从真实模拟的CIGRE中压网络的静态轨迹中研究线路选择性跳闸。一个卷积Q网络接收因果电压-电流相量及视在阻抗特征,可选地结合原始波形,并使用保守Q学习(CQL)进行训练。一项受控的敏感性研究在共同的数据划分和训练协议下评估了两种观测窗口、奖励变体以及三个CQL权重;一次探索性的事后运行额外将折扣因子从γ=0.95增加到0.99。在225个保留的情节中,最佳每时间步结果由组合输入和CQL权重α=0.9获得,达到了精确率0.9993、召回率0.9496和F1分数0.9738。由于密集的每时间步分数并未编码继电器操作的终端语义,我们还评估了每个情节中的第一个非等待动作。默认的组合输入智能体在214个故障情节中的98.13%中首先选择了正确的线路跳闸动作,但在11个非故障情节中的72.73%中发生了跳闸。在事后运行中,相应的比率分别为98.60%和54.55%。结果表明,密集预测性能和终端保护行为可能导致不同的模型排名。因此,离线CQL在模拟故障情节中展示了强大的故障线路选择能力,而静态轨迹、较小的非故障集和单一种子的事后设计排除了关于实际继电器安全性或部署准备度的结论。
英文摘要
Data-driven protection may complement conventional relays in distribution grids whose operating conditions vary with distributed generation, switching events, and changing short-circuit levels. We study line-selective tripping from static trajectories of a realistically simulated CIGRE medium-voltage network using offline reinforcement learning. A convolutional Q-network receives causal voltage-current phasor and apparent-impedance features, optionally together with raw waveforms, and is trained with conservative Q-learning (CQL). A controlled sensitivity study evaluates two observation windows, reward variants, and three CQL weights under a common split and training protocol; one exploratory post-hoc run additionally increases the discount factor from $γ$=0.95 to 0.99. On 225 held-out episodes, the best per-timestep result is obtained with combined input and CQL weight $α$=0.9, reaching precision 0.9993, recall 0.9496, and F1-score 0.9738. Because dense per-timestep scores do not encode the terminal semantics of relay operation, we also evaluate the first non-wait action in each episode. The default combined-input agent selects the correct line-trip action first in 98.13% of 214 fault episodes, but trips in 72.73% of the 11 non-fault episodes. In the post-hoc run, the corresponding rates are 98.60% and 54.55%, respectively. The results show that dense predictive performance and terminal protection behavior can lead to different model rankings. Offline CQL therefore demonstrates strong faulted-line selection on the simulated fault episodes, while the static trajectories, small non-fault set, and single-seed post-hoc design preclude conclusions about practical relay security or deployment readiness.
CommentsAccepted for presentation at the IEEE Power & Energy Student Summit (PESS 2026), Karlsruhe, Germany. 6 pages, 2 figures. Code: https://github.com/julianoelhaf/offline-cql-protection