发表机构
Illinois Institute of Technology; Meta Platforms; ByteDance/TikTok(伊利诺伊理工学院; Meta平台公司; 字节跳动/抖音)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本报告提出一种仅利用预测时信息的数据接入流程,检测并修复建筑计量数据缺陷,受控实验表明其能将梯度提升预测器误差恢复至清洁水平,并保留93%训练目标。
AI 中文摘要
电力公用事业公司和电网运营商日益依赖机器学习模型来预测次日需求,而这些模型所学习的数据经常存在缺陷:读数缺失、传感器冻结、建筑数小时读数为零,以及单位变化达100倍。本报告提出了一种数据接入流程,在模型训练前检测并修复此类缺陷,仅使用预测时可获得的信息,并通过受控实验衡量该流程对24小时前预测的保护效果。基于公共Building Data Genome 2数据集中12栋美国建筑的逐时电力数据(210,528行,2016-2017年),植入的、经哈希记录的缺陷影响训练期的0.10%,使梯度提升预测器的误差上升86%;经检测和仅基于过去数据的修复后,误差恢复到清洁数据水平(平均绝对缩放误差:清洁0.760,损坏1.415,修复0.729),同时保留了93%的训练目标。在根据已发表实地研究校准的缺陷发生率(训练行的1.6%)下,未受保护的预测器误差达到季节性朴素规则的4.4倍,而修复后的预测器再次匹配清洁基线。相同模式也适用于岭回归和随机森林,以及1至24小时的预测范围。流程的初始版本对自然数据过度清洁,使预测变差25%;该结果被保留,其原因在已发布工件中追溯,且修正此问题的逐建筑校准以带日期修正案的形式记录。每个数字均可通过固定的公共输入、SHA-256验证、84项自动化测试和持续集成复现。
英文摘要
Electric utilities and grid operators increasingly rely on machine-learning models to forecast next-day demand, and those models learn from meter data that is routinely defective: readings go missing, sensors freeze, buildings read zero for hours, and units change by a factor of 100. This report presents a data-onboarding pipeline that detects and repairs such defects before a model is trained, using only information available at forecast time, and a controlled experiment that measures whether the pipeline protects a 24-hour-ahead forecast. On hourly electricity data for twelve U.S. buildings from the public Building Data Genome 2 dataset (210,528 rows, 2016-2017), seeded, hash-logged defects touching 0.10% of the training period raised the error of a gradient-boosting forecaster by 86%; after detection and past-only repair the error returned to the clean-data level (mean absolute scaled error 0.760 clean, 1.415 corrupted, 0.729 repaired) while 93% of training targets were retained. At a defect prevalence calibrated to published field studies (1.6% of training rows) the unprotected forecaster's error reached 4.4 times that of a seasonal-naive rule, and the repaired forecaster again matched the clean baseline. The same pattern held for ridge regression and a random forest and across horizons of 1 to 24 hours. A first version of the pipeline over-cleaned natural data and made forecasts 25% worse; that result is retained, its cause is traced in the published artifacts, and the per-building calibration that corrects it is documented as a dated amendment. Every number is reproducible from pinned public inputs with SHA-256 verification, 84 automated tests and continuous integration.
CommentsTechnical report; 8 pages, 5 figures, 2 tables. Code, data manifests, and reproducibility artifacts: https://github.com/LIANGYIXUAN3335/energy-data-onboarding-forecasting . Preprint; not peer reviewed