arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02566cs.LG

通过分布式模型-智能体耦合在英国气象局统一模式中开展在线强化学习

Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling

Pritthijit Nath, Sebastian Schemm, Peter Haynes, Emily Shuckburgh, Mark Webb

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将英国气象局统一模式与分布式强化学习智能体耦合,在单案例中验证了该在线学习框架的可行性,可降低特定纬度带的气象预报误差,为业务系统应用奠定基础。

中文摘要 AI 辅助

机器学习校正方法需适配演变的模式状态,同时保持动力学一致性与数值稳定性,才能补充数值天气预报。为在全球预报模式中验证这一点,我们通过秩局部张量将英国气象局(UKMO)统一模式(UM)与分布式强化学习(RL)智能体耦合。深度确定性策略梯度(DDPG)的演员网络在每个大气柱的70个垂直模式层间共享权重,对模式倾向施加有界的位温校正。在10次有张弛(nudged)的训练预报中,向UKMO业务分析场进行张弛计算提供了即时反事实目标。随后将冻结策略在无张弛的预报中进行评估以用于推理。耦合工作流成功完成训练,且在评估案例中保持数值稳定。与匹配的原生UM +6小时预报相比,该学习策略在6个纬度带中的4个带降低了500百帕位势高度(Z₅₀₀)的平均绝对误差(MAE),其中北热带和南热带的降幅分别为45.8%和40.8%;海平面气压(MSLP)误差在3个带也有所降低,在0°-30°N带的最大降幅为27.3%。该单案例实验证明了分布式在线学习后接无张弛推理的显著潜力与可行性,为业务系统中基于RL的偏差校正和参数化奠定了基础。

英文摘要

Machine-learnt corrections can complement numerical weather prediction provided that they operate stably within an evolving numerical model. In this study, we couple the Met Office (UKMO) Unified Model (UM) with distributed reinforcement-learning agents through rank-local tensors. A column-aware deep deterministic policy gradient (DDPG) actor uses local vertical structure together with full-column context to apply bounded corrections to potential temperature and horizontal wind. During training, we perform ten nudged 6-hr 12-min forecasts, with nudging towards the UKMO operational analysis providing an immediate counterfactual target from which the policy learns. The resulting actor is then frozen and applied to a non-nudged forecast without access to analysis inputs or further weight updates. Relative to a matched non-nudged native forecast at +6 h, the corrected forecast reduces global latitude-weighted MAE by 2.85% for $Z_{500}$, 2.16% for MSLP, 5.16% for $T_{500}$ and 2.27% for $T_{1.5\textrm{m}}$, with an observed 3.57% wall-time overhead compared to native execution. Even though training and inference share the same initialisation, this single-case experiment demonstrates significant promise and feasibility, laying the groundwork for RL-based bias correction and parametrisations within operational systems.

发表机构

  • University of Cambridge(剑桥大学)
  • Met Office Hadley Centre(英国气象局哈德利中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑