发表机构
FEEIT Ss. Cyril and Methodius University in Skopje; Elektrodistribucija DOOEL EVN Group(斯科普里圣西里尔与美多德大学FEEIT学院; EVN集团下属Elektrodistribucija DOOEL公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对北马其顿17428台商用智能电表的两年用电数据,评估OWA、SoftImpute、形状建模自动编码器三种缺失值插补算法,发现OWA精度最高且稳定,SoftImpute稳定但精度低,自动编码器方差大,建议按需选算法并探索混合架构。
AI 中文摘要
通过高级计量基础设施(AMI)准确可靠地采集用电量数据对智能电网的运行至关重要,尤其对于非技术损失(NTL)的检测而言。然而,实际数据集常因通信故障出现缺失值。本文对三种用于大规模数据插补的高级算法进行实证评估:最优加权平均(OWA)方法、基于SoftImpute的低秩矩阵补全,以及形状建模自动编码器。现有关于用电量数据缺失值插补的研究往往缺乏在更大规模数据集上的验证。因此,本文的目标是在来自北马其顿的大规模实际用电量数据集上验证所选算法,该数据集包含17428台商用智能电表,覆盖两年时间。通过模拟数据中1至168小时的连续缺失区间,评估各算法的鲁棒性。结果表明,OWA在所有评估的缺失区间长度中提供最低的整体重构误差,且在长达一周的缺失区间的最坏场景中表现出极强的稳定性。相比之下,自动编码器表现出更高的方差,而SoftImpute虽稳定但精度较低。这些发现表明,应根据负荷曲线数据的特征选择插补方法,并强调未来电网管理系统中混合算法架构的潜力。
英文摘要
Accurate and reliable collection of electricity consumption data through Advanced Metering Infrastructure (AMI) is of great importance for the operation of smart grids, especially for the detection of non-technical losses (NTL). However, real-world datasets frequently suffer from missing values due to communication failures. This paper presents an empirical evaluation of three advanced algorithms for large-scale data imputation: the Optimally Weighted Average (OWA) method, Low-Rank Matrix Completion via SoftImpute, and a Shape-Modeling Autoencoder. Existing studies on missing value imputation in electricity consumption data often lack validation on larger datasets. Therefore, the goal of this paper is to validate the selected algorithms on a large-scale real-world electricity consumption dataset from North Macedonia that includes 17,428 commercial smart meters over two years. The robustness of each algorithm is evaluated by simulating continuous gaps in the data ranging from 1 to 168 hours. The results indicate that OWA provides the lowest overall reconstruction error across the evaluated gap sizes and strong stability in worst-case scenarios for gaps of up to one week. In contrast, the autoencoder exhibits higher variance, while SoftImpute has stable but inferior accuracy. These findings suggest that imputation methods should be selected based on the characteristics of load curve data and highlight the potential for hybrid algorithmic architectures in future grid management systems.
Comments6 pages, 7 figures; presented at the 2026 IEEE International Conference on Environment and Electrical Engineering and 2026 IEEE Industrial and Commercial Power Systems Europe (EEEIC / I&CPS Europe)