arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性下的太阳能光伏政策序列设计:一种基于智能体的强化学习方法

Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

Iias Faiud, Jonaid Shianifar, Michael Schukat, Karl Mason

arXiv 2609.04880首次发表:更新:

发表机构

School of Computer Science, University of Galway(戈尔韦大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将光伏政策设计建模为序列决策问题,结合强化学习与基于智能体模型,通过PPO等算法学习不同权衡的政策,其生成的政策权衡结构清晰且优于静态基准,可用于不确定性下的自适应政策设计。

AI 中文摘要

设计有效且财政可持续的太阳能光伏(PV)推广政策,需在不确定性与异质性决策背景下平衡推广收益与公共支出。本研究将光伏政策设计建模为序列决策问题,将强化学习(RL)与随机基于智能体模型(ABM)相结合,该模型可模拟不确定性下的年度太阳能光伏推广情况。政策制定者智能体在16年的时间范围内选择年度激励措施,包括资本补贴、优惠贷款利率和上网电价。研究在标量化奖励框架内通过调整政策偏好,探究推广与成本间的权衡关系。采用PPO、SAC和TD3算法学习政策,并在随机模拟环境中进行评估。结果表明,该方法形成了清晰的权衡结构:最高推广量政策(TD3,成本权重$w_{\text{cost}}=0.5$)实现约4145个推广者,成本为4173万欧元;最低成本政策(PPO,$w_{\text{cost}}=2.0$)将支出降至727万欧元,对应2682个推广者;平衡政策(PPO,$w_{\text{cost}}=1.6$)实现3495个推广者,成本为2247万欧元。不同算法均呈现一致的权衡模式,表明推广与成本关系具有鲁棒性,且RL框架相比静态基准政策能探索更广泛的政策配置。这些发现证明RL可作为不确定性下自适应政策设计的灵活工具。

英文摘要

Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent selects annual incentives, including capital grants, subsidised loan rates, and feed-in tariffs, over a 16-year horizon. Adoption--cost trade-offs are explored by varying policy preferences within a scalarised reward framework. Policies are learned using PPO, SAC, and TD3 and evaluated under stochastic simulation. The results show that this approach produces a clear trade-off structure: the highest-adoption policy (TD3, $w_{\text{cost}}=0.5$) achieves approximately 4,145 adopters at a cost of EUR 41.73 million, while the lowest-cost policy (PPO, $w_{\text{cost}}=2.0$) reduces expenditure to EUR 7.27 million with 2,682 adopters. The balanced policy (PPO, $w_{\text{cost}}=1.6$) achieves 3,495 adopters at a cost of EUR 22.47 million. Across algorithms, consistent trade-off patterns are observed, indicating robustness of the adoption--cost relationship. Compared with static baseline policies, the RL framework explores a broader range of policy configurations. These findings demonstrate the potential of RL as a flexible tool for adaptive policy design under uncertainty.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑