基于执行的评估揭示语言模型在环境科学计算中的隐藏失败
Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science
浏览论文内容
中文总结 AI 辅助
该研究构建了环境科学计算基准AtmosCoder-Bench,发现现有评估未观测计算过程,多项选择格式会虚高准确率,模型多因未一致应用公式失败,前沿模型仍需专家监督。
中文摘要 AI 辅助
大型语言模型正越来越多地用于环境科学的定量工作,但现有评估仅对最终答案打分,未观测计算过程。本文提出AtmosCoder-Bench,这是一个基于执行的基准,可使计算过程可视化。该基准通过可迁移的半自动化流水线构建,包含436个问题、3910个变体、7029个已分级的量,每个问题都经过验证,具有明确性、人类可解性,且答案可唯一验证。研究发现:(i)多项选择格式使测得的准确率至少提高了12个百分点;(ii)许多失败并非源于知识缺失,而是模型在多步骤计算中未能始终如一地应用已知公式和约束;(iii)即使是前沿模型,在特定任务条件使熟悉的方法失效时仍表现薄弱,常回归到规范的解决方案模式,而非使方法适应相关物理机制,因此专家监督至关重要。
英文摘要
Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human-solvable, with uniquely verifiable answers. We find that (i) multiple-choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi-step computation; and (iii) even frontier models remain weak when task-specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.
发表机构
- Department of Geography, Hong Kong(香港地理系)
机构由 AI 辅助整理,请以论文原文为准。