AI 中文总结
研究电网问题中图神经网络单任务微调的拓扑过拟合故障,提出MxGPS多路图变换器,通过自监督预训练和多任务微调联合训练SSE与PF,实验表明其能有效防止拓扑过拟合,以少得多参数实现拓扑无关泛化。
AI 中文摘要
针对电网问题的图神经网络单任务微调存在系统故障模式:在分布内误差最低的模型在拓扑变化时退化最严重。我们将此称为拓扑过拟合,即特定任务梯度信号倾向于编码训练拓扑特有的关系结构而非潜在物理特性。为揭示并解决此故障模式,我们引入了MxGPS,它通过自监督预训练和多任务微调协议,在共享节点编码器上运行K个任务专用的GPS分支,联合训练静态状态估计(SSE)和交流潮流(PF),并通过消融实验评估跨分支注意力模块。联合SSE+PF目标迫使共享编码器同时满足互补梯度信号,防止其过度拟合特定拓扑的关系结构。在跨越四个未见拓扑(14、24、162和300节点)的3折滑动窗口交叉验证中,MxGPS在所有四个零样本潮流拓扑上的边界违规率(BVR)为0%。关键的是,分布内PF误差低得多的模型在拓扑变化时退化190%至1400%,而MxGPS仅退化39%,这直接表明拓扑过拟合是故障机制而非模型能力不足。MxGPS仅160万个参数(比GridFM参考基线少12倍),证明了多任务联合训练是电网基础模型中拓扑无关泛化的一种有原则且参数高效的机制。
英文摘要
Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting: the tendency of task-specific gradient signals to encode relational structure particular to the training topologies rather than the underlying physics, causing models to fail on unseen grids despite strong in-distribution performance. To expose and address this failure mode, we introduce MxGPS (Multiplex GPS), a multiplex graph transformer that runs K task-specialised GPS branches over a shared node encoder, jointly trained on Static State Estimation (SSE) and AC Power Flow (PF) via a self-supervised pre-training and multi-task fine-tuning protocol, with a cross-branch attention module evaluated in ablation. The joint SSE+PF objective forces the shared encoder to simultaneously satisfy complementary gradient signals, preventing it from overfitting to topology-specific relational structure. Under a 3-fold sliding-window cross-validation spanning four unseen topologies (14-, 24-, 162-, and 300-bus), MxGPS attains 0% boundary violation rate (BVR) on all four zero-shot Power Flow topologies. Critically, models with substantially lower in-distribution PF error degrade by 190% to 1400% under topology shift, whereas MxGPS degrades by only 39%, an inversion that directly implicates topology overfitting as the failure mechanism rather than insufficient model capacity. With only 1.6M parameters (12x fewer than the GridFM reference baseline), MxGPS demonstrates that multi-task joint training is a principled and parameter-efficient mechanism for topology-agnostic generalisation in power grid foundation models.
Comments10 pages, 4 figues