arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过数据缩放推动高分辨率天气预报的极限

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun

arXiv 2608.14652首次发表:更新:

发表机构

State Key Laboratory of Severe Weather Meteorological Science and Technology, Nanjing University; School of Atmospheric Science, Nanjing University; Ant Healthcare, Ant Group; DAMO Academy, Alibaba Group(南京大学灾害性天气气象科技国家重点实验室; 南京大学大气科学学院; 蚂蚁集团蚂蚁医疗; 阿里巴巴集团达摩院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对高分辨率天气预报的 data bottleneck,提出BaguanHR框架,通过变量级超分辨率合成高分辨率数据,使预报性能超越现有方法,且数据量与预报误差呈幂律关系。

AI 中文摘要

基于机器学习(ML)的0.1°全球天气预报模型的发展受到高分辨率数据可用性有限的制约,因为数十年的再分析数据仅以0.25°的分辨率提供。虽然现有方法在有限的0.1°样本上对0.25°预报模型进行微调,但我们表明,这种迁移受到粗分辨率预报固有的不可逆信息丢失的阻碍。因此,我们提出BaguanHR,一个将重点从迁移模型转向迁移数据的框架。我们首先表明,超分辨率(SR)比预报具有更低的条件熵和输入放大,使其成为分辨率迁移更稳健的载体。通过利用这一优势,我们通过变量级超分辨率从ERA5合成了大量0.1°数据。BaguanHR在合成加真实数据集上的性能超过了基于ML的方法和IFS-HRES,在72小时内超过85%的提前期实现了更优性能。此外,我们的发现突出了幂律缩放效应,数据量翻倍可使72小时预报的RMSE降低4.6%,120小时预报的RMSE降低4.9%。我们的结果表明,缩放基于ML的高分辨率预报主要是数据瓶颈,而变量级超分辨率提供了一种简单且通用的解决方案,以解锁长期粗分辨率再分析数据用于高分辨率训练。

英文摘要

The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25$^{\circ}$ resolution. While existing approaches fine-tune 0.25$^{\circ}$ forecast models on limited 0.1$^{\circ}$ samples, we show that this transfer is hindered by the irreversible information loss inherent in coarse-resolution forecasting. Therefore, we propose BaguanHR, a framework that shifts the focus from transferring models to transferring data. We first show that super-resolution (SR) has lower conditional entropy and input amplification than forecasting, making it a more robust vehicle for resolution transfer. By leveraging this advantage through variable-wise SR, we synthesize extensive 0.1$^{\circ}$ data from ERA5. BaguanHR's performance on the synthetic-plus-real dataset exceeds both ML-based methods and IFS-HRES, achieving superior performance across over 85% of the lead times within 72 hours. Furthermore, our findings highlight a power-law scaling effect, as a twofold increase in data reduces RMSE by 4.6% for 72-hour forecasting and 4.9% for 120-hour forecasting. Our results demonstrate that scaling high resolution ML-based forecasting is primarily a data bottleneck, and that variable-wise super-resolution provides a simple yet general solution to unlock long coarse-resolution reanalyses for high-resolution training.

CommentsAccepted by ECCV2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑