每个固定度量都有盲点:一种用于评分预报真实感的学习型大气评判器
Every Fixed Metric Has a Blind Spot: A Learned Atmospheric Critic for Scoring Forecast Realism
- ETH Zürich(苏黎世联邦理工学院)
- ETH AI Center, ETH Zürich(苏黎世联邦理工学院人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对机器学习天气预报模型的伪影问题,提出学习型判别器作为大气评判器,以自适应检测失败模式并评分预报真实感,实验表明优于现有度量且能区分模型类型。
AI中文摘要:
尽管机器学习天气预报模型在逐点度量上具有高精度,但它们可能表现出不同的失败模式,如模糊、周期性不规则以及其他非物理空间伪影。这促使人们提出了多种度量来检测已知的失败案例。现有度量预先固定了表示或变换,而这种选择限制了它们能检测到的伪影类型。我们提出训练一个判别器,用于将参考数据与模型的输出分开,并利用其输出对数几率(logit)获得一种类似散度的真实感评分。该判别器学习区分模型场与真实天气的任何特征,并适应模型所表现出的任何失败模式。我们将我们的学习型大气评判器与现有度量进行比较,使用应用于ERA5再分析数据的各种合成损坏。我们的方法成功识别了这些损坏并对其严重性进行排序,而现有度量至少在一个损坏上失败。此外,我们评估了真实天气模型的预报,发现真实感评分随着预报时效的延长而降低,并且该度量通常赋予数值模型比机器学习模型更高的真实感。
英文摘要:
Despite their high accuracy on point-wise metrics, machine learning weather forecasting models can exhibit different failure modes such as blurring, periodic irregularities, and other unphysical spatial artifacts. This has motivated a variety of metrics to detect known failure cases. Existing metrics fix a representation or transformation in advance, and that choice limits the artifacts they can detect. We propose to train a discriminator for separating reference data from the model's output, and using its output logit to obtain a divergence-like realism score. The discriminator learns whatever separates the model's fields from real weather, adapting to whichever failure mode that model exhibits. We compare our learned atmospheric critic to existing metrics using various synthetic corruptions applied to ERA5 reanalysis data. Our method successfully identifies the corruptions and ranks their severity, while existing metrics fail on at least one corruption. Additionally, we evaluate forecasts from real weather models, and find that the realism score degrades with longer lead times and the metric generally assigns higher realism to numerical models than to machine learning models.