发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究首次在真实站点观测上对生成式天气数据同化进行受控基准测试,发现生成先验优于3D-Var,全梯度引导最有效,扩散与流匹配性能相当,并建立了标准化基准。
AI 中文摘要
天气再分析产品依赖于计算密集的数值天气预报,随后通过数据同化将预报向观测修正。深度生成模型提供了一种更廉价的替代方案,将大部分成本从推理阶段转移到离线训练阶段。然而,现有的生成方法已在合成观测或不同数据集和评估方案下进行评估,使得不清楚哪些设计选择实际上改善了真实世界的数据同化。我们提出了首个在真实天气站点观测上对生成式天气数据同化进行受控基准测试。使用美国本土11,849个NOAA MADIS站点和四个天气变量,我们在保持数据集、观测算子和深度学习架构固定的情况下评估方法。该基准比较了主要设计选择,包括扩散与流匹配、像素与潜空间公式,以及多种推理时条件化策略,并与经典3D-Var基线进行对比。基准揭示了三个明确结论。第一,学习到的生成先验优于3D-Var的高斯先验(相对于ERA5的RMSE降低35.7% vs. 33.3%),尽管在推理时不使用ERA5背景场。第二,全梯度引导始终优于停止梯度和初始噪声优化。第三,其他选择几乎没有可衡量的益处:在匹配条件下,扩散和流匹配表现几乎相同,潜空间变量混合也无帮助。我们进一步评估了密集和稀疏站点设置,发现生成式AI和全梯度引导的优势在稀疏条件下更为显著。这些结果共同确定了生成式天气数据同化中哪些组件能提高真实站点观测的性能,并为未来工作建立了标准化基准。
英文摘要
Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Recent advances in deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices improve real-world data assimilation. We present the first controlled benchmark of generative data assimilation for single-time near-surface analysis from real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four near-surface variables, we hold the dataset, observation operator, and deep learning architecture fixed, and measure spatial generalization at held-out stations. The benchmark compares the major design choices proposed for generative data assimilation, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three conclusions. First, the best generative methods outperform 3D-Var (35.7% vs. 33.3% RMSE improvement over ERA5), although 3D-Var receives the ERA5 field at the analysis time as its background and the generative methods receive none. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. Both advantages widen when stations are sparse. Together, these results identify which components of generative data assimilation improve spatial generalization in near-surface analysis.