arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化比较预测模型评估中协议引发的不确定性:来自大规模每日PM10预报的证据

Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting

Rafael da Silva, Kiersten Monahan

arXiv 2609.26288首次发表:更新:

发表机构

Eastern University(东方大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究量化了预测模型评估中协议切换引发的排名不确定性,提出PSS框架,通过大规模PM10预报实验证明协议间位移远超协议内扰动,建议报告排名稳定性。

AI 中文摘要

预测模型的比较研究通常以对候选模型进行排名告终,然而这些排名依赖于评估协议,而协议的影响很少被视为不确定性来源。我们将此问题形式化为协议引发的排名不确定性,并引入一个框架,该框架比较因切换协议引起的排名位移与在固定协议内因常规选择引起的位移。我们使用协议敏感性得分(PSS)和全重拟合协议内参考来量化这些效应。我们在每日PM10的大规模序列预测研究中验证了该框架。静态分割和滚动原点评估在425个欧洲背景站和365个美国环保署监测站上进行比较。切换协议产生的平均PSS值分别为0.801和0.772,并分别在35.3%和31.5%的站点改变了所选模型。在欧洲,具有相同评分目标的协议内扰动产生的PSS值为0.072和0.230,胜者交换率分别为0.8%和4.9%。因此,协议间位移显著大于所选协议内参考。将候选集从三个模型扩展到九个模型,使欧洲的协议间胜者交换率增加到60.2%。该模式在应用于保留的背景和非背景站点的冻结协议下也持续存在。这些结果表明,模型选择的结论可能实质性依赖于合理的评估选择。我们建议在声称模型优越性时,报告在一小组可辩护的协议内扰动下的排名稳定性。

英文摘要

Comparative studies of predictive models often end by ranking candidate models, yet these rankings depend on evaluation protocols whose influence is rarely treated as a source of uncertainty. We formalize this problem as protocol-induced ranking uncertainty and introduce a framework that compares ranking displacement caused by switching protocols with displacement produced by conventional choices within a fixed protocol. We quantify these effects using the Protocol Sensitivity Score (PSS) and a full-refit intraprotocol reference. We validate the framework in a large-scale sequential prediction study of daily PM10. Static-split and rolling-origin evaluation are compared across 425 European background stations and 365 US EPA monitors. Switching protocols produces mean PSS values of 0.801 and 0.772 and changes the selected model at 35.3% and 31.5% of stations, respectively. In Europe, intraprotocol perturbations with identical scored targets produce PSS values of 0.072 and 0.230, with winner-swap rates of 0.8% and 4.9%. Between-protocol displacement is therefore substantially larger than the selected within-protocol references. Expanding the candidate set from three to nine models increases the between-protocol winner-swap rate to 60.2% in Europe. The pattern also persists under a frozen protocol applied to held-out background and non-background stations. These results show that model-selection conclusions can depend materially on legitimate evaluation choices. We recommend reporting ranking stability under a small set of defensible intraprotocol perturbations alongside claims of model superiority.

Comments29 pages, 8 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑