商业电池调度在分时电价下的MILP、MPC与强化学习比较评估
Comparative Evaluation of MILP, MPC, and Reinforcement Learning for Commercial Battery Dispatch Under Time-of-Use Tariffs
浏览论文内容
中文总结 AI 辅助
本研究基于2023年商业光伏数据,比较MILP、MPC和SAC强化学习在分时电价下的电池调度性能,发现MPC能恢复MILP基准99.2%的经济效益,而SAC表现不佳。
中文摘要 AI 辅助
电池储能系统(BESS)与屋顶光伏(PV)配套使用,在分时(TOU)电价下可带来可衡量的成本节约;然而,基于模型和无模型的调度策略在全年、真实商业数据集上的相对性能仍未得到充分基准测试。本文利用一个在分时电价下运行的商业光伏电站数据,对三种BESS调度方法进行了全年(2023年)比较评估。所考察的策略包括:(i)具有完美预见性的混合整数线性规划(MILP)公式,在所假设模型下提供基准性能(oracle benchmark);(ii)基于日前持久性预测的模型预测控制(MPC)方案,代表一种低复杂度的可部署方法;以及(iii)在因果信息约束下训练的软演员-评论家(SAC)深度强化学习智能体。MILP基准相对于无储能基线实现了24.6%的年成本降低。基于持久性的MPC方法仅利用前一日数据即可恢复该基准的99.2%。相比之下,所评估的SAC智能体产生的年成本高于无储能基线。这一结果在能源系统强化学习已知挑战的背景下进行了分析,包括有限可观测性和奖励设计。总体而言,结果表明,对于所研究的数据集和电价结构,基于持久性的MPC在实际部署约束下捕获了几乎所有可实现的经济效益,而所考虑的RL配置在相同信息限制下未能产生有竞争力的性能。
英文摘要
Battery energy storage systems (BESS) paired with rooftop photovoltaics (PVs) can deliver measurable cost savings under time-of-use (TOU) electricity tariffs; however, the relative performance of model-based and model-free dispatch strategies remains insufficiently benchmarked on full-year, real-world commercial datasets. This paper presents a full-year (2023) comparative evaluation of three BESS dispatch approaches using data from a commercial PV installation operating under a TOU tariff. The examined strategies include: (i) a mixed-integer linear programming (MILP) formulation with perfect foresight, providing an oracle performance benchmark under the assumed model; (ii) a model predictive control (MPC) scheme based on a day-ahead persistence forecast, representing a low-complexity deployable approach; and (iii) a soft actor-critic (SAC) deep reinforcement learning agent trained under causal information constraints. The MILP benchmark achieves an annual cost reduction of 24.6\% relative to a no-storage baseline. The persistence-based MPC approach recovers 99.2\% of this benchmark using only prior-day data. In contrast, the evaluated SAC agent yields an annual cost higher than the no-storage baseline. This outcome is analyzed in the context of known challenges in reinforcement learning for energy systems, including limited observability and reward design. Overall, the results indicate that, for the studied dataset and tariff structure, persistence-based MPC captures nearly all achievable economic benefits under practical deployment constraints, whereas the considered RL configuration does not yield competitive performance under the same information limitations.
发表机构
- Lappeenranta-Lahti University of Technology(拉彭兰塔-拉赫蒂理工大学)
- University of Vaasa(瓦萨大学)
- Cleanwatts Digital S.A(Cleanwatts Digital 股份公司)
机构由 AI 辅助整理,请以论文原文为准。