arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据泄露夸大了停电预测模型的可推广性

Data Leakage Inflates Generalizability of Power Outage Prediction Models

Yamil Essus, Ranga Raju Vatsavai, Benjamin Rachunok

arXiv 2608.24665首次发表:更新:

发表机构

University of Toronto; North Carolina State University(多伦多大学; 北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出停电预测模型因数据泄露夸大了可推广性,评估时采用更贴近现实的空间/时间/事件拆分后性能大幅下降,GeoAI嵌入仅小幅改善空间泛化,需改进数据与评估方式。

AI 中文摘要

停电预测模型越来越多地用于气候驱动的基础设施风险评估,但当前的评估实践模糊了这些模型是否能推广到此类应用所需的新场景。我们确定了停电预测模型中三种常见的方法选择,这些选择会影响模型在空间、时间和事件相关场景下的泛化能力。我们使用2018年至2023年美国东海岸的公开数据、来自天气再分析和土地覆盖数据的特征集,以及来自GeoAI基础模型Prithvi WxC的嵌入,比较了不同方法决策对预测性能的影响。具体而言,我们在多种测试选择策略下评估模型性能,包括未过滤的随机拆分、留一州拆分和留一事件拆分设计,这些设计越来越接近现实世界的部署条件。虽然随机训练-测试拆分产生了良好的性能,但我们表明这些结果因空间和时间自相关而被夸大。在空间和时间留存实验中,预测准确性大幅下降,模型常常无法优于简单的空基准。纳入GeoAI基础模型嵌入带来的改进有限且不一致,主要体现在空间泛化方面,且无法解决事件层面的迁移性差的问题。这些发现表明,鉴于当前的数据可用性和评估实践,公开训练的停电预测模型提供的操作价值有限且不确定。进展可能需要改进数据覆盖范围、更现实的评估协议,以及将重点从边际模型改进转向解决结构性数据约束。

英文摘要

Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require. We identify three common methodological choices in power outage prediction models that influence their ability to generalize across spatial, temporal, and event-based settings. We compare the predictive performance impacts of different methodological decisions using publicly available data for the U.S. East Coast from 2018 to 2023 and feature sets derived from weather reanalysis and land-cover data, and embeddings from a GeoAI foundation model (Prithvi WxC). Specifically, we assess model performance under multiple test selection strategies, including unfiltered random splits, leave-one-state-out, and leave-one-event-out designs, which increasingly approximate real-world deployment conditions. While random train-test splits yield strong performance, we show that these results are inflated by spatial and temporal autocorrelation. Under spatial and temporal holdout experiments, predictive accuracy degrades substantially, with models often failing to outperform a simple null baseline. Incorporating GeoAI foundation model embeddings yields limited and inconsistent improvements, primarily for spatial generalization, and does not resolve poor event-level transferability. These findings suggest that, given current data availability and evaluation practices, publicly trained outage prediction models offer limited and uncertain operational value. Progress will likely require improved data coverage, more realistic evaluation protocols, and a shift in focus from marginal modeling advances toward addressing structural data constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑