arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NestyNet. I. 物理函数难以用神经网络拟合:一种用于精确替代模型及解析导数的框架

NestyNet. I. Physics Functions Are Hard to Fit with Neural Networks: A Framework for Accurate Surrogates and Analytic Derivatives

Rodrigo Ibata, Wassim Tenachi, Foivos Diakogiannis, Neil Ibata, Anirudh Shankar

arXiv 2608.05862首次发表:更新:

AI 中文总结

NestyNet是耦合模型与优化器的框架,可高精度拟合物理平滑函数并输出解析导数,在AI Feynman基准测试中性能远超标准神经网络,速度也显著快于自动微分基线。

AI 中文摘要

物理学中许多至关重要的平滑函数,恰恰是标准神经网络方法难以精确拟合的对象。本文提出NestyNet,这是一种耦合模型与优化器的框架,能够高精度拟合此类目标函数,同时以低成本解析方式输出其梯度、海森矩阵、拉普拉斯算子及不定积分。这使其成为科学机器学习任务的天然基础。该模型是确定性分段解析替代模型,其优化器为二阶Levenberg--Marquardt方案,该方案的阻尼项与线性求解器针对物理及其他科学应用中典型的多尺度、尖锐结构目标所诱导的刚性、强相关参数几何进行了定制。在包含120个物理方程的AI Feynman基准测试中,相较于采用一阶优化(Adam)训练的标准神经网络,NestyNet在函数值上实现了2100倍的中位数提升,一阶导数提升1400倍,二阶导数提升780倍;即便在采用拟牛顿法(L-BFGS)优化对拟合结果进行精细化后,相应的提升仍达540倍、450倍和250倍。由于其解析设计,它比向量化自动微分(autograd)基线快约44倍,且优势随模型规模增大而增长。该解析导数框架还支持向量值与复值目标、输入和输出的测量不确定性以及约束条件,其模块可灵活组合以构建科学上有用的模型架构,全程无需依赖autograd。这些组件共同构成了一个实用的模块化框架,用于拟合复杂的科学替代模型,同时为后续分析提供精确的微分算子。

英文摘要

Many of the smooth functions that matter most in physics are precisely the ones that standard neural network methods struggle to fit accurately. Here we present NestyNet, a coupled model-and-optimizer framework capable of fitting such targets to high accuracy while also delivering their gradients, Hessians, Laplacians, and antiderivatives analytically and at low cost. This makes it a natural substrate for scientific machine learning tasks. The model is a deterministic segmented analytic surrogate, and its optimizer is a second-order Levenberg--Marquardt scheme whose damping and linear solves are tailored to the stiff, strongly correlated parameter geometries induced by multiscale and sharply structured targets typical in physics and other scientific applications. On the AI Feynman benchmark of 120 physics equations, NestyNet achieves median improvement factors of $2\,100\times$ for function values, $1\,400\times$ for first derivatives, and $780\times$ for second derivatives relative to standard neural networks trained with first-order optimization (Adam). Even after refining those fits with quasi-Newton (L-BFGS) optimization, the corresponding improvements are $540\times$, $450\times$, and $250\times$. Owing to the analytic design it is up to $\approx 44\times$ faster than vectorized automatic-differentiation (autograd) baselines, with a margin growing with model size. The same analytic-derivative framework also supports vector- and complex-valued targets, measurement uncertainties in both inputs and outputs, and constraints, and its modules can be composed flexibly to build scientifically useful model architectures, all without reverting to autograd. Together, these components provide a practical modular framework for fitting difficult scientific surrogates while delivering accurate differential operators for subsequent analysis.

Comments25 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑