arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SS-ESOAP:面向物理信息学习的自缩放自适应预处理

SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke, Yixuan Wang, Anima Anandkumar

arXiv 2608.29448首次发表:更新:

发表机构

McGill University; California Institute of Technology; Aarhus University; University of Copenhagen(麦吉尔大学; 加州理工学院; 奥胡斯大学; 哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对物理信息神经网络的病态目标问题,提出SS-ESOAP优化方法,在8个PDE基准中多数取得最低残差,在Boussinesq问题上优于Adam,为刚性高精度物理信息训练提供可扩展选项。

AI 中文摘要

物理信息神经网络(PINNs)常面临病态目标问题,限制了高精度训练。稠密拟牛顿方法可改善局部条件,但需要昂贵的优化器状态;而Kronecker分解方法如SOAP可扩展到更大的网络,但依赖周期性基更新。我们提出\textit{\textbackslash method},它在SOAP风格预处理的基础上,增加了适配Kronecker几何的标量割线能量修正,以及自适应基更新后的方差状态下缩放。我们刻画了标量修正诱导的定向割线匹配,并给出了基变换过程中方差状态失配的界。在8个PDE基准测试中,\textit{\textbackslash method}在6个基准上取得了最低的最终残差,包括Burgers和Boussinesq问题,而SOAP系列基线在Gray-Scott和Ginzburg-Landau问题上表现更好。在Boussinesq问题上,\textit{\textbackslash method}在4.1小时内达到了$10^{-5}$的残差,峰值VRAM为9.2 GB,而Adam在14小时内未达到该目标。在4个代表性PDE上的三种子集$L^2$和$H^1$误差,支持了更低残差与更高解精度之间的关联。这些结果表明,\textit{\textbackslash method}是针对刚性、高精度物理信息训练的可扩展选项,而非现有优化器的统一替代品。

英文摘要

Physics-informed neural networks (PINNs) often face ill-conditioned objectives that limit high-accuracy training. Dense quasi-Newton methods improve local conditioning but require expensive optimizer state, while Kronecker-factored methods such as SOAP scale to larger networks but rely on periodic basis updates. We introduce \method, which augments SOAP-style preconditioning with a scalar secant-energy correction adapted to Kronecker geometry and an adaptive basis update followed by variance-state downscaling. We characterize the directional secant matching induced by the scalar correction and give a bound on variance-state mismatch across basis changes. Across eight PDE benchmarks, \method attains the lowest final residual on six, including Burgers and Boussinesq, while SOAP-family baselines perform better on Gray-Scott and Ginzburg-Landau. On Boussinesq, \method reaches a residual of $10^{-5}$ in 4.1 hours with 9.2 GB peak VRAM, while Adam does not reach this target within 14 hours. Three-seed $L^2$ and $H^1$ errors on four representative PDEs support the link between lower residuals and improved solution accuracy. These results position \method as a scalable option for stiff, high-accuracy physics-informed training, rather than a uniform replacement for existing optimizers.

Comments31 pages, 12 figures, 16 tables. Submitted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑