发表机构
State Key Laboratory of Severe Weather Meteorological Science and Technology, Nanjing University; School of Atmospheric Science, Nanjing University; Ant Healthcare, Ant Group; DAMO Academy, Alibaba Group(南京大学灾害性天气气象科技国家重点实验室; 南京大学大气科学学院; 蚂蚁集团蚂蚁医疗; 阿里巴巴集团达摩院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对高分辨率天气预报的 data bottleneck,提出BaguanHR框架,通过变量级超分辨率合成高分辨率数据,使预报性能超越现有方法,且数据量与预报误差呈幂律关系。
AI 中文摘要
基于机器学习(ML)的0.1°全球天气预报模型的发展受到高分辨率数据可用性有限的制约,因为数十年的再分析数据仅以0.25°的分辨率提供。虽然现有方法在有限的0.1°样本上对0.25°预报模型进行微调,但我们表明,这种迁移受到粗分辨率预报固有的不可逆信息丢失的阻碍。因此,我们提出BaguanHR,一个将重点从迁移模型转向迁移数据的框架。我们首先表明,超分辨率(SR)比预报具有更低的条件熵和输入放大,使其成为分辨率迁移更稳健的载体。通过利用这一优势,我们通过变量级超分辨率从ERA5合成了大量0.1°数据。BaguanHR在合成加真实数据集上的性能超过了基于ML的方法和IFS-HRES,在72小时内超过85%的提前期实现了更优性能。此外,我们的发现突出了幂律缩放效应,数据量翻倍可使72小时预报的RMSE降低4.6%,120小时预报的RMSE降低4.9%。我们的结果表明,缩放基于ML的高分辨率预报主要是数据瓶颈,而变量级超分辨率提供了一种简单且通用的解决方案,以解锁长期粗分辨率再分析数据用于高分辨率训练。
英文摘要
The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25$^{\circ}$ resolution. While existing approaches fine-tune 0.25$^{\circ}$ forecast models on limited 0.1$^{\circ}$ samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1$^{\circ}$ data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.
CommentsAccepted by ECCV2026