发表机构
Stanford University; University of Southern California; University of Utah(斯坦福大学; 南加州大学; 犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大气化学微调的人工智能基础模型学到了什么,通过对微软Aurora模型施加化学扰动、检查内部表示等方法,发现其虽捕捉到部分臭氧响应但未执行化学约束,提供了测试模型是否学大气化学的框架,强调预测应依内部机制评判。
AI 中文摘要
天气预报基础模型(FMs)越来越多地被微调以预测空气质量,能以比传统化学传输模型更低的计算成本提供快速的全球污染预测。这些FMs通常基于再分析数据训练并通过自回归展开生成预测,未明确表示物理或化学过程。本文通过研究微软的Aurora模型,首次探讨了针对大气化学微调的FM学到了什么。对其预测施加受控化学扰动并根据已知光化学关系进行测试,然后检查生成这些预测的内部表示。发现Aurora捕捉到了对活性氮的一阶臭氧响应,但未执行基于过程的模型所编码的化学约束。其内部表示在很大程度上仍围绕预训练期间继承的气象学组织,几乎没有特定于化学的结构。使用稀疏自动编码器识别出因果控制化学预测但未清晰映射到单个大气过程的内部组件。这项工作提供了一个测试人工智能预测系统是否从再分析数据中学到大气化学的框架。随着这些模型越来越多地用于为环境政策决策提供信息,我们认为成分预测也应根据其内部机制而非仅根据基准技能来判断。
英文摘要
Weather forecasting foundation models (FMs) are increasingly fine-tuned to predict air quality, offering fast global pollution forecasts at lower computational cost than conventional chemical transport models. These FMs are typically trained on reanalysis data and generate forecasts through autoregressive rollout. They do not explicitly represent governing physical or chemical processes. Therefore, high forecast skill does not reveal whether a model has learned physical mechanisms or exploits statistical regularities in its training data. Here, we present the first study of what a FM fine-tuned for atmospheric chemistry has learned by examining Microsoft's Aurora model. We impose controlled chemical perturbations on its forecasts and test them against known photochemical relationships. We then examine the internal representations that generate these forecasts. We find that Aurora captures a first-order ozone response to reactive nitrogen but does not enforce the chemical constraints that a process-based model encodes. It generates chemically inconsistent combinations of related species and relaxes localized emission features such as wildfire plumes toward background. Internally, its representations remain largely organized around the meteorology inherited during pretraining, with little structure specific to chemistry. Using sparse autoencoders, we identify internal components that causally control the chemical forecast but do not map cleanly onto individual atmospheric processes. This work provides a framework for testing whether AI forecasting systems learn atmospheric chemistry from reanalysis data. As these models are increasingly positioned to inform environmental policy decisions, we argue that composition forecasts should also be judged by their internal mechanisms rather than by benchmark skill alone.