arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于重放的输电扩展协同优化的学习稳定性研究:考虑策略性投标

Learning Stability of Replay-Based Co-Optimization for Transmission Expansion under Strategic Bidding

Tomonari Kanazawa, Hikaru Hoshino, Eiko Furutani

arXiv 2610.09366首次发表:更新:

发表机构

University of Hyogo(兵库县立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究策略性投标下输电扩展协同优化中重放记忆大小与学习率的权衡,提出最近邻滤波排除不一致样本,以提升稳定性并避免振荡。

AI 中文摘要

本文研究了电力市场中策略性投标下基于学习的输电扩展协同优化的行为。在该框架中,输电容量被更新,同时市场参与者通过深度强化学习学习其投标策略,导致耦合且非平稳的学习动态。我们表明,投标智能体的瞬态策略退化可能产生不一致的成本-容量样本,这些样本会偏置输电容量更新,并阻止协同优化过程收敛到期望解。为缓解这一问题,引入了带有最近邻滤波的重放式容量更新,以排除更新数据中的不一致样本。随后,我们分析了当重放记忆扩大时出现的一种新的振荡行为。在IEEE 30节点系统上的数值结果表明,扩大的重放记忆提高了对瞬态策略退化的鲁棒性,但可能引入时间滞后,导致振荡,除非适当降低容量更新的学习率。这些结果揭示了重放记忆大小与学习率之间的权衡,并为稳定协同优化提供了实用指导。

英文摘要

This paper investigates the behavior of learning-based co-optimization for transmission expansion under strategic bidding in electricity markets. In this framework, transmission capacities are updated while market participants simultaneously learn their bidding strategies through deep reinforcement learning, resulting in coupled and non-stationary learning dynamics. We show that transient policy degradation of bidding agents can generate inconsistent cost-capacity samples, which bias the transmission-capacity update and prevent the co-optimization process from converging to the desired solution. To mitigate this, replay-based capacity updates with nearest-neighbor filtering are introduced to exclude inconsistent samples from the update data. We then analyze a new oscillatory behavior that appears when the replay memory is enlarged. Numerical results on the IEEE 30-bus system demonstrate that enlarged replay memories improve robustness against transient policy degradation but can introduce a temporal lag, leading to oscillations unless the capacity-update learning rate is appropriately reduced. These results reveal the trade-off between replay-memory size and learning rate and provide practical guidelines for stable co-optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑