当AI用于偏微分方程(PDE)时,何时能产生科学证据?
When Does AI for PDEs Yield Scientific Evidence?
浏览论文内容
中文总结 AI 辅助
该研究针对AI用于PDE领域中,现有基准仅评估模型预测精度、未评估证据支持的问题,扩展相关基准并揭示精度与证据支持的排名差异,明确了评估与应用的不匹配情况。
中文摘要 AI 辅助
现有的AI用于偏微分方程(PDE)的基准主要根据预测或近似精度评估模型。然而在物理学研究中,AI输出常作为科学主张的证据。这两个目标并不等价:前者衡量输出与参考目标的一致性或对控制约束的满足度;后者要求,在给定研究对象、科学主张、假设和证据标准的情况下,输出是否为该主张提供足够证据。为弥合这一差距,我们扩展了广泛使用的PDE模拟基准和全面的PDE反问题基准,首次在AI用于PDE的领域中,评估模型输出是否支持指定的科学主张以及支持的程度。我们的结果显示,数值精度和证据支持对模型的排名可能不同,解释了这种差异出现的时间和原因,并揭示现有基准可能倾向于那些输出对目标科学主张提供较弱支持的方法。综上,我们在AI用于PDE的领域中,对这种评估与应用不匹配的情况进行了形式化、实证演示和解释。
英文摘要
Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measures an output's agreement with a reference target or satisfaction of governing constraints; the latter asks whether, given a specified object of study, scientific claim, assumptions, and evidence standard, the output provides sufficient evidence for that claim. To bridge this gap, we extend a widely used PDE-simulation benchmark and a comprehensive benchmark for PDE inverse problems to enable, for the first time in AI for PDEs, evaluation of whether and to what extent model outputs support specified scientific claims. Our results show that numerical accuracy and evidential support can rank models differently, explain when and why they do so, and reveal that existing benchmarks can favor methods whose outputs provide weaker support for the scientific claims of interest. Together, we formalize, empirically demonstrate, and explain this evaluation--use mismatch in AI for PDEs.
发表机构
- School of Future Technology, South China University of Technology(华南理工大学未来技术学院)
机构由 AI 辅助整理,请以论文原文为准。