HetGPS:面向电动汽车充电的、带物理锚定自适应安全机制的可扩展图多智能体强化学习
HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive Safety for EV Charging
浏览论文内容
中文总结 AI 辅助
该研究提出HetGPS混合图控制框架,结合学习图风险与物理锚定校正,实现可扩展的电动汽车充电多智能体协调,降低电压违规率并提升离网成功率,模型规模与车队无关且可跨系统零样本迁移。
中文摘要 AI 辅助
针对网络耦合智能体大群体的安全干预,需在保护共享约束的同时避免过度覆盖面向任务的策略决策。本文提出HetGPS,一种混合图控制框架,通过将干预幅度与校正方向分离,结合学习到的图风险与物理锚定校正:动作条件图残差模型调度依赖状态的干预权限,物理模型确定干预方向。针对电动汽车(EV)充电,将该滤波器与参数共享的异构图软Actor-Critic(SAC)策略耦合,实现感知拓扑的协调,且学习到的模型规模与车队规模无关。在5个嵌套配电网(含200--3218辆EV,100个评估日)中,自适应权限机制将母线-步电压违规率从无滤波时的3.93--7.74%降至0.52--3.44%,同时维持99.06--100%的离网成功率;与相同物理定向投影且权限固定的方法相比,其在所有5个网络上的平均奖励均提升,且在4个网络上的平均安全评分降低。部署的策略与风险模型在所有规模下均含383,702个学习参数;在3218辆EV场景中,匹配的集中式SAC演员规模约为其170倍。在8变压器系统上训练的策略可零样本迁移至16和32变压器系统,违规率达0.57--0.75%,离网成功率至少99.99%。上述结果表明,学习到的图风险可大规模分配干预权限,而馈线物理机制锚定校正动作。
英文摘要
Safety interventions for large populations of network-coupled agents must protect shared constraints without unnecessarily overriding task-oriented policy decisions. We present HetGPS, a hybrid graph-control framework synergizing learned graph risk with physics-anchored correction by separating intervention magnitude from corrective direction. An action-conditioned graph residual model schedules state-dependent intervention authority, while a physics model determines its direction. For electric vehicle (EV) charging, we couple this filter with a parameter-shared heterogeneous graph soft actor-critic policy, enabling topology-aware coordination with a learned model size independent of fleet size. Across five nested distribution networks with 200--3,218 EVs and 100 evaluation days, Adaptive Authority reduces bus--step voltage violations from 3.93--7.74\% without filtering to 0.52--3.44\%, while maintaining 99.06--100\% departure success. Relative to the same physics-directed projection with fixed authority, it improves mean reward on all five networks and lowers the mean safety score on four. The deployed policy-and-risk model contains 383,702 learned parameters at every scale; at 3,218 EVs, a matched centralized SAC actor is about $170\times$ larger. A policy trained on the eight-transformer system transfers zero-shot to the 16- and 32-transformer systems, attaining 0.57--0.75\% violation rates and at least 99.99\% departure success. These results show that learned graph risk can allocate intervention authority at scale while feeder physics anchors corrective action.