arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.12198physics.ao-phcs.AIcs.LG

探索基于Skew-T图的多模态AI推理用于气象预报

Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams

  • Korea Meteorological Administration(韩国气象厅)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Chungnam National University(忠南大学)
  • AIPIM

机构由 AI 辅助整理,请以论文原文为准。

ChangJae Lee, Heecheol Yang, Jonghak Choi

更新

AI总结:

本文提出一个轻量级多模态AI助手,通过课程学习和思维链推理解读Skew-T图以估计降水概率,其性能可比肩业务NWP模型,且计算高效、可解释性强。

AI中文摘要:

基于大气探空资料进行预报是业务气象学中的一项基础任务,通常需要人类预报员对Skew-T log-P图进行结构化的视觉推理。尽管视觉语言模型(VLMs)的最新进展已在其他科学领域展现出潜力,但其在气象图解读中的应用仍基本未被探索。在本研究中,我们提出了一个轻量级AI助手,利用一个小型语言模型(LM)和一个经微调以模拟人类预报员的小型VLM来解读Skew-T图。采用课程学习框架,我们首先通过视觉问答训练模型从图中识别关键大气特征,随后进行思维链推理任务,基于所得到的视觉定位估计降水概率。模型输入包括文本摘要或由业务数值天气预报(NWP)预报生成的Skew-T图,并与韩国自动气象站网络的三小时降水观测数据配对。评估结果表明,尽管仅依赖静态大气廓线,微调后的VLM仍达到了与业务NWP模型相当的技巧。消融研究表明,视觉定位和推理监督对性能至关重要,而注意力图分析证实模型学会关注相关的气象特征。这些发现凸显了紧凑、可解释的多模态模型在支持天气预报任务方面的潜力。该方法为大规模系统提供了一种计算高效的替代方案,未来工作可将其扩展到更复杂的应用。

英文摘要:

Forecasting from atmospheric soundings is a fundamental task in operational meteorology, often requiring structured visual reasoning over Skew-T log-P diagrams by human forecasters. While recent advances in Vision-Language Models (VLMs) have shown promise in other scientific domains, their application to meteorological diagram interpretation remains largely unexplored. In this study, we present a lightweight AI assistant that interprets Skew-T diagrams using a small language model (LM) and a small VLM fine-tuned to emulate human forecasters. Using a curriculum learning framework, we first train the models to identify key atmospheric features from diagrams through visual question answering, followed by chain-of-thought reasoning tasks that estimate precipitation probability based on the derived visual groundings. Model inputs include either textual summaries or generated Skew-T diagrams derived from operational Numerical Weather Prediction (NWP) forecasts, paired with three-hour precipitation observations from South Korea's Auto Weather Stations network. Evaluation results demonstrate that the fine-tuned VLM achieves skill comparable to an operational NWP model, despite relying solely on static atmospheric profiles. Ablation studies reveal that visual grounding and reasoning supervision are critical for performance, while attention map analysis confirms that the model learns to focus on relevant meteorological features. These findings highlight the potential of compact, interpretable multimodal models to support weather forecasting tasks. The approach offers a computationally efficient alternative to large-scale systems, and future work could extend it to more complex applications.

补充信息

↑