随机扰动权重:从确定性机器学习天气模型构建集成
Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models
浏览论文内容
中文总结 AI 辅助
本文提出随机扰动权重(SPW)方法,在推理时扰动确定性机器学习天气模型的权重以生成集成,无需重新训练,在10天预报时效下达到接近训练概率模型的技能,但噪声注入位置需按架构调参。
中文摘要 AI 辅助
机器学习天气模型(MLWMs)在全球中期预报中现已达到或超越业务数值天气预报(NWP)的性能,且推理成本远低于后者。许多已部署的MLWMs是确定性的,仅产生单一预报,无法估计其自身的不确定性;而日益增多的训练概率模型则直接生成校准的集成,但代价是需要专门的训练过程。我们转而探讨:从已存在的确定性检查点中,在不重新训练的情况下,能提取出多少不确定性?物理集成通过随机扰动参数化倾向来表示模型不确定性,我们则在推理时扰动网络的原始权重张量,这一方案称为随机扰动权重(SPW)。我们还探究该方法是否有效、应在何处及何种尺度上注入噪声,以及它在何处失效。一项跨越四个确定性骨干模型(Aurora、GraphCast、SFNO和AIFS)的三阶段消融研究,为每个模型选定一个生产基线,并与训练概率模型AIFS-ENS、FourCastNet 3和Atlas以及业务ECMWF集成(IFS-ENS)在112个初始化时刻上进行对比。在240小时(10天)预报时效下,SPW集成的连续排序概率技能评分(CRPSS)比最佳训练概率基线低0.04至0.13,但边际训练成本为零。没有一种噪声注入位置适用于所有模型:有效的张量组因架构而异,因此SPW目前是一种调参流程,而非即插即用的通用方案。其主要失效模式是产生一致的全场偏移,导致区域均值过度离散;将噪声限制在粗尺度或扰动初始条件可分别部分修复该问题。
英文摘要
Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run. We ask instead how much uncertainty can be extracted from a deterministic checkpoint that already exists, without retraining it. Where physical ensembles represent model uncertainty by stochastically perturbing parametrisation tendencies, we perturb the network's raw weight tensors at inference time, a scheme we call stochastically perturbed weights (SPW). We also ask whether it works, where and on which scales to inject the noise, and where it fails. A three-phase ablation across four deterministic backbones, Aurora, GraphCast, SFNO, and AIFS, selects one production baseline per model, benchmarked against the trained-probabilistic AIFS-ENS, FourCastNet 3 and Atlas as well as the operational ECMWF ensemble (IFS-ENS) over 112 initialisation times. At a 240 h (10-day) lead time the SPW ensembles reach continuous ranked probability skill scores (CRPSS) between 0.04 and 0.13 below the best trained-probabilistic baseline, at zero marginal training cost. No injection site works across models: the productive tensor group is architecture-specific, so SPW is at present a tuning procedure rather than a plug-and-play recipe. Its main failure mode is a coherent whole-field offset that overdisperses the domain mean, and restricting the noise to coarse scales or perturbing the initial conditions each repair part of it.
发表机构
- Federal Office for Meteorology and Climatology MeteoSwiss(瑞士联邦气象与气候局)
- ETH Zurich(苏黎世联邦理工学院)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。