arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

地球科学中AI模型的战略治理

Strategic Governance of AI Models in Earth Science

Makoto Kelp, Amirhossein Arzani, Patricia Castellanos, Paul Griffiths, Ivan Higuera-Mendieta, Manuel Perez-Carrasco, Viral Shah, Patrick Obin Sturm, James Weber

arXiv 2610.10560首次发表:更新:

发表机构

University of Utah; NASA Goddard Space Flight Center; University of Bristol; Stanford University; Center for Astrophysics | Harvard & Smithsonian; GESTAR II, NASA Global Modeling and Assimilation Office; University of Southern California; University of Reading(犹他大学; 美国国家航空航天局戈达德太空飞行中心; 布里斯托大学; 斯坦福大学; 哈佛-史密森天体物理中心; GESTAR II,美国国家航空航天局全球建模与同化办公室; 南加州大学; 雷丁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文指出地球科学AI模型的预测技能与物理可靠性存在差异,确定了其物理评估的五项优先事项,并建议未来十年开展开放AI评估数据集等三项工作以加强战略治理。

AI 中文摘要

在天气和气候数据上预训练的AI基础模型,正越来越多地被微调用于远超天气预报范畴的地球科学任务。它们的开发与应用速度已超过科学界对其进行评估的能力。这些模型几乎完全通过基准技能指标来评判,该指标衡量的是预测结果与参考产品的吻合程度,却未考量模型是否能体现其所预测系统的物理过程。因此,预测技能与物理可靠性是两种截然不同的属性。在这些模型从未接受过训练的气候变化非平稳条件下,这种区分最为关键。我们针对地球科学领域从特定任务模拟器到基础模型的AI模型的物理评估,确定了五项优先事项,涵盖训练数据、微调、行为测试、机制可解释性以及输出验证。我们建议在未来十年开展三项工作:1)开放可用于AI评估的数据集;2)建立基于物理评估的共享报告标准;3)设立专门针对这些模型安全性的研究计划。

英文摘要

AI foundation models pretrained on weather and climate data are increasingly fine-tuned to Earth science tasks well beyond weather forecasting. Their development and adoption are outpacing the scientific community's ability to evaluate them. These models are judged almost entirely by benchmark skill metrics, which measure how closely a forecast reproduces a reference product but not whether a model represents the physical processes governing the system it predicts. Forecast skill and physical reliability are therefore distinct properties. The distinction is most consequential under the nonstationary conditions of a changing climate for which these models were never trained. We identify five priorities for the physical evaluation of AI models in Earth science from task-specific emulators to foundation models, spanning training data, fine-tuning, behavioral testing, mechanistic interpretability, and output validation. We recommend three activities for the coming decade: 1) open AI-ready evaluation datasets, 2) a shared reporting standard for physics-based evaluation, and 3) a dedicated research program on the safety of these models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑