arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciExam for ENSO:AI 智能体能否构建气候模型?

SciExam for ENSO: Can AI Agents Build Climate Models?

Yinling Zhang, Langchen Liu, Dongbin Xiu, Xueyan Zou, Xu Kuang, Mengdi Wang, Shilong Liu

arXiv 2610.10513首次发表:更新:

发表机构

The Ohio State University; Yale University; University of California, San Diego; Stanford University; Princeton University(俄亥俄州立大学; 耶鲁大学; 加利福尼亚大学圣迭戈分校; 斯坦福大学; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SciExam for ENSO基准,让AI智能体基于真实观测在六小时内构建ENSO低阶随机模型,并通过隐藏评分器评估。十二个系统中六个超越已发表模型,较强模型结构与冷暖不对称性的公开争论相容,表明智能体能在无已知答案下构建有竞争力的科学模型。

AI 中文摘要

语言模型智能体越来越多地被要求开展开放式的科学研究,然而其结果通常是根据已知答案、评分标准或语言模型评审员来评分的,这些方式都无法判断一个新的科学模型是否有效。针对厄尔尼诺-南方涛动现象的AI科学考试(SciExam for ENSO)是一个基准测试,在该测试中,智能体根据真实观测数据构建厄尔尼诺-南方涛动(ENSO)的低阶随机模型,ENSO是年际气候变率的主导模态。在六小时的预算内,智能体处理观测数据,编写自己的诊断程序(这些诊断程序随后被冻结),并仅利用这些诊断作为反馈来开发模型。隐藏的评分者随后测试该模型是否能重现ENSO的统计特征、恢复未观测变量、预测留出的年份,并以相同方式对已发表的模型进行评分。在十二个智能体系统中,有六个生成的模型得分高于已发表模型,主要归功于更好的重建和预测能力。较强模型的简化形式分别与关于ENSO冷暖不对称性的两种相互竞争的解释之一相容,而这一公开争论在任务中从未提及。在多种信息条件下对顶级系统的受控运行表明,其得分并非来自对过时观测记录的回忆,而且其接收到的信息塑造了其构建模型的方式。因此,SciExam for ENSO可以在没有已知答案的情况下评估智能体的研究,结果表明智能体已经能够构建具有竞争力的模型,其结构涉及科学家仍在争论的问题。

英文摘要

Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.

Comments28 pages, 5 figures, 8 tables. Code: https://github.com/ylzhang2447/SciExam-ENSO-code

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑