发表机构
College of Computer Science and Technology, Jilin University; School of Big Data and Artificial Intelligence, Guangdong University of Finance and Economics; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University; School of Computer Science, Zhuhai College of Science and Technology; Department of Mechanical Engineering, National University of Singapore(吉林大学计算机科学与技术学院; 广东财经大学大数据与人工智能学院; 吉林大学符号计算与知识工程教育部重点实验室; 珠海科技学院计算机科学学院; 新加坡国立大学机械工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在从噪声数据发现偏微分方程,提出LLM-PDESR框架,结合大语言模型假设生成与数学评估环境,利用五次样条和子域加权残差减轻噪声影响,经帕累托驱动反馈环优化方程,实验表明其在多方面显著优于现有方法。
AI 中文摘要
从噪声观测数据中发现控制偏微分方程(PDEs)是科学机器学习中的一项基本挑战。传统符号回归(SR)方法在庞大组合搜索空间中难以识别准确方程,因其无法纳入特定领域先验知识,且依赖逐点评估和离散有限差分会放大高频噪声。为此提出LLM-PDESR框架,将大语言模型的结构假设生成与严格数学评估环境相结合。采用C^4连续五次样条进行稳健微分,子域加权残差作为低通滤波器,减轻了困扰现有方法的适应度景观失真。通过帕累托驱动反馈环使大语言模型迭代优化候选方程。在23个经典PDEs和5个新方程上评估LLM-PDESR,它能从噪声ERA5再分析数据中成功提取一维动态代理(1D-CACE)的一致结构框架,在结构恢复、抗噪声能力等方面显著优于现有方法。
英文摘要
Discovering governing partial differential equations (PDEs) from noisy observational data is a fundamental challenge in scientific machine learning. Traditional symbolic regression (SR) methods often struggle to identify accurate equations within vast combinatorial search spaces, largely due to their inability to incorporate essential domain-specific prior knowledge. Furthermore, reliance on pointwise evaluations and discrete finite differences inherently amplifies high-frequency noise, creating deceptive fitness landscapes that derail the optimization process. To resolve these bottlenecks, we propose LLM-PDESR, a framework that integrates the structural hypothesis generation of Large Language Models (LLMs) with a mathematically rigorous evaluation environment. By employing C^4-continuous quintic splines for robust differentiation and subdomain weighted residuals as natural low-pass filters, our approach effectively mitigates the fitness landscape distortion that plagues existing methods. A Pareto-driven feedback loop then enables the LLM to iteratively refine candidate equations, balancing predictive accuracy with structural parsimony. We evaluate LLM-PDESR on 23 canonical PDEs and five structurally novel equations (including a multivariate system) specifically designed to preclude dataset memorization and test true discovery capabilities. Demonstrating real-world applicability, the framework successfully extracts a consistent structural skeleton for an interpretable 1D dynamical surrogate (1D-CACE) directly from noisy ERA5 reanalysis data. Extensive experiments and out-of-distribution testing confirm that LLM-PDESR significantly outperforms state-of-the-art methodologies in structural recovery, noise resilience, and the avoidance of spurious complexity and equation bloat.
Comments28 pages, 12 figures