arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

风电场控制的离线强化学习:动态风向下的风洞实验研究

Offline Reinforcement Learning for Wind Farm Control: A Wind Tunnel Study under Dynamic Wind Directions

Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao

arXiv 2609.12905首次发表:更新:

发表机构

University of Warwick; Technical University of Munich(华威大学; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对动态风向下的风电场功率最大化问题,提出离线强化学习算法MTD3-BC,通过偏航控制减轻尾流效应,风洞实验表明其比贪婪策略提升约10%功率,且无需尾流模型、训练成本低,为首次实验验证。

AI 中文摘要

本文研究了风向变化情境下的风电场功率最大化问题。具体而言,提出了一种无模型的、结合行为克隆的改进型双延迟深度确定性策略梯度(MTD3-BC)算法,通过偏航控制在变化的风向条件下解决该任务。MTD3-BC是一种离线强化学习(RL)算法,旨在仅从预先收集的离线数据集中推断出良好行为。此外,为确保偏航调整的平滑与适度,在策略优化目标中引入了一项新的动作一致性项。与在线RL方法不同,MTD3-BC在训练期间无需与风电场模拟器进行大量交互,显著降低了计算成本和训练时间。为验证算法在变化风向下的有效性,开展了一项风洞实验。结果表明,MTD3-BC成功减轻了尾流效应,相较于基线贪婪策略,实现了约10%的场级功率增益,其性能与基于数据校准的模型尾流转向基准相当,且无需尾流模型,训练成本仅为在线RL的一小部分。据我们所知,这是离线RL风电场控制策略首次得到实验验证与展示。

英文摘要

This paper addresses the wind farm power maximization problem in the presence of wind direction changes. Specifically, a model-free Modified Twin Delayed Deep Deterministic Policy Gradient with Behavior Cloning (MTD3-BC) algorithm is proposed to tackle this task through yaw control under varying wind direction conditions. MTD3-BC is an offline reinforcement learning (RL) algorithm that aims to infer good behavior from only a precollected offline dataset. Additionally, to ensure smooth and moderate yaw adjustments, a new action consistency term is introduced into the policy optimization objective. Unlike online RL methods, MTD3-BC does not require extensive interactions with a wind farm simulator during training, significantly reducing computational costs and training time. A wind tunnel experiment is conducted to validate the effectiveness of the algorithm under varying wind directions. The results demonstrate that MTD3-BC successfully mitigates wake effects, delivering farm-level power gains of approximately 10\% over the baseline greedy strategy and performance on par with a data-calibrated model-based wake-steering benchmark, while requiring no wake model and only a small fraction of the training cost of online RL. To our knowledge, this is the first time an offline RL wind farm control policy has been validated and demonstrated experimentally.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑