发表机构
University of Warwick; Technical University of Munich(华威大学; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态风向下的风电场功率最大化问题,提出离线强化学习算法MTD3-BC,通过偏航控制减轻尾流效应,风洞实验表明其比贪婪策略提升约10%功率,且无需尾流模型、训练成本低,为首次实验验证。
AI 中文摘要
本文研究了风向变化情境下的风电场功率最大化问题。具体而言,提出了一种无模型的、结合行为克隆的改进型双延迟深度确定性策略梯度(MTD3-BC)算法,通过偏航控制在变化的风向条件下解决该任务。MTD3-BC是一种离线强化学习(RL)算法,旨在仅从预先收集的离线数据集中推断出良好行为。此外,为确保偏航调整的平滑与适度,在策略优化目标中引入了一项新的动作一致性项。与在线RL方法不同,MTD3-BC在训练期间无需与风电场模拟器进行大量交互,显著降低了计算成本和训练时间。为验证算法在变化风向下的有效性,开展了一项风洞实验。结果表明,MTD3-BC成功减轻了尾流效应,相较于基线贪婪策略,实现了约10%的场级功率增益,其性能与基于数据校准的模型尾流转向基准相当,且无需尾流模型,训练成本仅为在线RL的一小部分。据我们所知,这是离线RL风电场控制策略首次得到实验验证与展示。
英文摘要
This paper addresses the wind farm power maximization problem in the presence of wind direction changes. Specifically, a model-free Modified Twin Delayed Deep Deterministic Policy Gradient with Behavior Cloning (MTD3-BC) algorithm is proposed to tackle this task through yaw control under varying wind direction conditions. MTD3-BC is an offline reinforcement learning (RL) algorithm that aims to infer good behavior from only a precollected offline dataset. Additionally, to ensure smooth and moderate yaw adjustments, a new action consistency term is introduced into the policy optimization objective. Unlike online RL methods, MTD3-BC does not require extensive interactions with a wind farm simulator during training, significantly reducing computational costs and training time. A wind tunnel experiment is conducted to validate the effectiveness of the algorithm under varying wind directions. The results demonstrate that MTD3-BC successfully mitigates wake effects, delivering farm-level power gains of approximately 10\% over the baseline greedy strategy and performance on par with a data-calibrated model-based wake-steering benchmark, while requiring no wake model and only a small fraction of the training cost of online RL. To our knowledge, this is the first time an offline RL wind farm control policy has been validated and demonstrated experimentally.