arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向时空车辆到电网调度的带动态边界的分层约束强化学习

Hierarchical Constrained Reinforcement Learning with Dynamic Boundary for Spatio-Temporal Vehicle-to-Grid Scheduling

Haoyu Yan, Shutong Ding, Jiebao Zhang, Xi Yao, Yu Liu, Haoyu Wang, Chenchi Luo, Ye Shi

arXiv 2608.03409首次发表:更新:

AI 中文总结

针对大规模电动汽车V2G调度的时空约束与计算瓶颈问题,提出HPC-RL分层框架,经实验其性能优于基线方法,计算时间大幅缩短,约束满足性优异。

AI 中文摘要

电动汽车(EV)的快速普及给电网带来了显著的时空不确定性,而车辆到电网(V2G)技术通过双向功率流动提供了关键的灵活性。然而,将大规模电动汽车整合到最优潮流框架中面临巨大挑战,原因在于求解器复杂性带来的计算瓶颈以及耦合的时空约束。现有的强化学习(RL)方法在高度动态的电动汽车车队环境中,往往难以在严格的约束满足性与可扩展性之间取得平衡。为应对这些挑战,本文提出了一种用于时空耦合V2G调度的分层约束强化学习策略(HPC-RL)框架。该框架采用双层架构:上层利用基于广义约梯度法的RL算法严格执行电网级空间硬约束;下层实现一种新型动态边界策略,为单个电动汽车计算实时可行的充电功率边界,从而确保时间充电需求得到满足。这种集成设计不仅能在RL优化过程中同时处理时空耦合约束,还能通过分层解耦显著提升大规模车队的泛化能力。在IEEE 14、30以及改进的141节点系统上开展的大量实验表明,HPC-RL在所有指标上均优于模型预测控制及最先进的安全RL基线。所提方法实现了近最优的调度策略,在大规模场景中将在线计算时间从数小时大幅缩短至数分钟,同时保持近零的约束违反率和近100%的充电需求满足率。

英文摘要

The rapid proliferation of Electric Vehicles (EVs) introduces significant spatio-temporal uncertainties into power grids, while Vehicle-to-Grid (V2G) technology offers critical flexibility through bidirectional power flow. However, integrating large-scale EVs into the Optimal Power Flow framework presents substantial challenges due to computational bottlenecks arising from solver complexity and coupled spatio-temporal constraints. Existing Reinforcement Learning (RL) methods often struggle to balance strict constraint satisfaction with scalability in highly dynamic EV fleet environments. To address these challenges, this paper proposes a Hierarchical Policy for Constrained Reinforcement Learning (HPC-RL) framework for spatially and temporally coupled V2G scheduling. The framework adopts a two-layer architecture: the upper level utilizes a RL algorithm based on the Generalized Reduced Gradient method to strictly enforce spatial grid-level hard constraints; the lower level implements a novel dynamic boundary strategy to compute real-time feasible charging power bounds for individual EVs, thereby ensuring the satisfaction of temporal charging demands. This integrated design not only enables the simultaneous handling of spatially and temporally coupled constraints during the RL optimization process but also significantly enhances generalization capabilities for large-scale fleets through hierarchical decoupling. Extensive experiments on IEEE 14, 30, and modified 141-bus systems demonstrate that HPC-RL outperforms Model Predictive Control and state-of-the-art safe RL baselines across all metrics. The proposed method achieves near-optimal scheduling strategies and drastically reduces online computation time in large-scale scenarios from hours to minutes, while maintaining a near-zero constraint violation rate and nearly 100\% charging demand satisfaction.

CommentsWAICA 2026 Best Student Paper Award Runner-Up

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑