数值神经算子:数值分析与算子学习的关联及从有限元方法构建离散神经算子
Numerical Neural Operator: Connection Between Numerical Analysis and the Operator Learning and Building Discretized Neural Operators from Finite Element Methods
- University of Notre Dame(圣母大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探究数值分析与神经算子的关联,从有限元等数值格式推导神经算子架构,建立收敛误差估计,其网络规模与近似误差呈多项式依赖,实验验证了该方法的低误差与高参数效率。
AI中文摘要:
算子学习旨在学习函数空间之间的映射,已被广泛应用于与偏微分方程(PDE)相关的问题,这引发了两个基本问题:第一,鉴于标准数值方法和神经算子均旨在近似PDE解算子,二者之间是否存在关联?第二,当神经算子在数值生成的数据上进行训练时,它是近似潜在的无限维算子、训练数据本身,还是用于生成训练数据的离散数值格式?本文的关键假设是,训练后的神经算子可能会模拟生成训练数据的离散数值格式,该假设已通过Chen Chen 1995年论文中提出的线性近似原理及后续的尺度定律论文得到部分初步验证,但从数值离散化角度开展的神经算子近似研究仍未被探索。在本研究中,我们考虑一类PDE并从数值格式中推导神经算子架构:首先,数值基函数(如有限元基函数)可通过可学习神经网络近似或直接构建;其次,用于时间离散化的有限差分格式可展开为迭代神经架构,以生成所得神经算子基的系数。通过利用数值分析,我们为这些从数值推导而来的神经算子建立了收敛误差估计,值得注意的是,网络规模在L₂范数下与近似误差呈多项式依赖关系,这改进了一般神经算子近似中推导的嵌套指数依赖关系,并与数值观测结果一致。最后,数值实验表明,该方法可降低误差并提高参数效率。
英文摘要:
Operator learning aims to learn mappings between function spaces and has been widely used for PDE-related problems. This raises two fundamental questions. First, is there a connection between standard numerical methods and neural operators, given that both are designed to approximate PDE solution operators? Second, when a neural operator is trained on numerically generated data, does it approximate the underlying infinite dimension operator, the training data themselves, or the discretized numerical scheme used to generate them? A key hypothesis in the paper is that the trained neural operator may mimic the discretized numerical schemes that generate the training data. The hypothesis has been partially and preliminary verified by the linear approximation principle proposed in Chen Chen 1995 paper and the following scaling law papers. However, neural operator approximation from the perspective of numerical discretization remains unexplored. In this work, we consider a class of PDEs and derive neural operator architectures from the numerical schemes. Firstly, the numerical basis functions, such as finite element basis functions, can be approximated or directly constructed by learnable neural networks. Then the finite difference schemes for temporal discretization can be unrolled into iterative neural architectures that generate the coefficients of the basis of the resulting neural operator. By leveraging numerical analysis, we establish convergence error estimates for these numerically derived neural operators. Notably, the network size depends polynomial on the approximation error in $L_2$, improving the nested exponential dependence derived in the general neural operator approximation and aligning with the numerical observations. Finally, numerical experiments demonstrate reduced errors and improved parameter efficiency.