发表机构
University of Michigan-Dearborn; Pacific Northwest National Laboratory(密歇根大学迪尔伯恩分校; 太平洋西北国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出混合两阶段机器学习流水线,解耦故障检测与分类,在输电故障检测中显著提升准确率,优于基准方法,解决类别不平衡及特征模糊问题。
AI 中文摘要
高压输电网络中快速准确的故障检测对电网可靠性和设备保护至关重要。输电故障数据集常存在类别不平衡问题,某些故障类型产生的电气特征落在正常运行范围内,导致单模型分类器在安全关键场景中失效。本文提出一种混合两阶段机器学习流水线,将检测与分类解耦。第一阶段结合孤立森林(Isolation Forest)异常检测器与可选的监督二分类器,采用或融合规则;监督分支在训练时自动分配给异常检测器无法解决的任何故障类别,当不存在此类类别时则省略。第二阶段仅对第一阶段标记的样本应用随机森林(Random Forest)多分类器。特征工程被表述为每个测量点的算子,将6个原始通道映射为18个特征,包括基于Fortescue定理推导的零序对称分量,因此L个测量点对应18L个特征。在TLFaultDataset数据集上,该流水线将线路故障端到端准确率从31.3%提升至95.8%;在独立单点数据集上,同一框架在所有类别(含正常运行)上达到97.25%的端到端准确率,无需GPU或联邦基础设施,且在CPU上每个样本仅需0.05毫秒,超过TLFed联邦基准的94.84%。对两个数据集的 ablation 实验显示,零序特征可解决三相与三相接地的模糊性,使该类别对的F1分数从0.39提升至0.997。研究还发现,零序特征的方向与系统相关,因此需采用学习得到的决策边界而非固定继电器阈值。
英文摘要
Rapid and accurate fault detection in high-voltage transmission networks is essential for grid reliability and equipment protection. Transmission fault datasets are frequently imbalanced, and certain fault types produce electrical signatures that fall within the normal operating envelope, causing single-model classifiers to fail on safety-critical cases. This paper proposes a hybrid two-stage machine learning pipeline that decouples detection from classification. Stage 1 combines an Isolation Forest anomaly detector with an optional supervised binary detector through an OR-fusion rule; the supervised branch is allocated automatically during training for any fault class the anomaly detector cannot resolve, and is omitted when no such class exists. Stage 2 applies a Random Forest multiclass classifier only to samples flagged by Stage 1. Feature engineering is expressed as a per-measurement-point operator mapping six raw channels to eighteen features, including zero-sequence symmetrical components derived from Fortescue's theorem, yielding 18L features for L measurement points. On the TLFaultDataset, the pipeline raises Line-fault end-to-end accuracy from 31.3% to 95.8%. On an independent single-point dataset, the same framework attains 97.25% end-to-end accuracy across all classes including normal operation, exceeding the TLFed federated benchmark of 94.84% without GPU or federated infrastructure, at 0.05 ms per sample on CPU. Ablation on both datasets shows zero-sequence features resolving the three-phase versus three-phase-to-ground ambiguity, raising the F1-score of that class pair from 0.39 to 0.997. The direction of the zero-sequence signature is found to be system-dependent, motivating a learned decision boundary in place of a fixed relay threshold.
Comments21 pages