用于多变量风电场SCADA数据过滤的聚类算法
Clustering algorithms for multivariate wind farm SCADA data filtering
浏览论文内容
中文总结 AI 辅助
研究风电场SCADA数据过滤问题,通过比较多种聚类算法与手动过滤的精度,引入适用评估指标,应用于某海上风电场数据,结果显示多数聚类方法精度更高,但不同模型有差异,专家参与仍必要。
中文摘要 AI 辅助
在风电场运行期间,监控与数据采集(SCADA)系统记录了大量异常、瞬变和特定运行模式,产生了庞大的数据集。但众多应用仅需正常运行的测量值,所以SCADA数据必须过滤。为此已提出多种方法来自动化并取代专家通过目视检查数据进行的手动过滤。本文比较了多种聚类算法与手动过滤的过滤精度,引入适用于未标记数据且在潜在应用中稳健的评估指标。基于结果,给出了将模型校准推广到不同数据集的建议并讨论各模型的潜在用例。将模型应用于现有海上风电场三台涡轮机的SCADA数据,使用多个数据通道上的10分钟统计数据。除了通常记录的异常和运行模式,该数据集因多次现场测试存在大量不明显的异常值。总体而言,结果突出了在特征选择和评估指标设计中扩展超出功率曲线分析的重要性。多数情况下,基于聚类的方法能检测明显和细微的异常值,比手动过滤精度更高。然而,精度和保留的数据量因模型而异,专家参与仍有必要,不过相比手动过滤程度降低。
英文摘要
During wind farm operation, Supervisory Control and Data Acquisition (SCADA) systems record numerous anomalies, transients, and specific operational modes, leading to large datasets. However, for a wide range of applications, only measurements corresponding to normal operation are required and, therefore, the SCADA data must be filtered. For this purpose, several methods have been proposed to automate and replace manual filtering conducted by experts via visual inspection of the data. In this paper, we compare the filtering accuracy of multiple clustering algorithms against manual filtering, introducing evaluation metrics that are suitable for unlabeled data and robust across potential applications. Based on the results, we provide recommendations for generalizing model calibration to different datasets and discuss potential use cases for each model. The models are applied to the SCADA data of three turbines of an existing offshore wind farm, using 10-minute statistics across multiple data channels. In addition to the anomalies and operational modes typically recorded, the dataset presents a large number of non-evident outliers due to several field tests. Overall, the results highlight the importance of extending the analysis beyond the power curve, both in feature selection and in the design of evaluation metrics. In most cases, cluster-based methods are able to detect both evident and subtle outliers, achieving higher accuracy than manual filtering. However, the accuracy and the amount of data retained vary considerably depending on the model, and expert involvement remains necessary, though to a reduced extent compared to manual filtering.