arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21352math.AP

用于非均匀耦合抛物系统的变分物理信息神经网络训练动力学中的双重筛选

Double screening in the training dynamics of variational physics-informed neural networks for heterogeneous coupled parabolic systems

Ali Ouattara Kouma, Gossrin Jean-Marc Bomisso

首次发表
浏览论文内容

中文总结 AI 辅助

研究非均匀耦合抛物系统的变分物理信息神经网络训练动力学,通过双重筛选定理得出四个结果,包括耦合能量恒等式等,指出条件数增长影响共享步长梯度下降,Adam优化器可减轻困难,且预测经数值验证。

中文摘要 AI 辅助

我们分析了应用于线性耦合抛物对流 - 扩散 - 反应系统的变分物理信息神经网络的训练动力学,当只有部分组件经历对流传输时称为非均匀系统。在神经切线核区域,时空变分残差上的梯度流简化为线性微分系统,其算子是由系统算子的时空符号和矩阵切线核构建的Gram矩阵。主要结果是双重筛选定理。在主导对流下,该矩阵相对于对流组件块的舒尔补收敛到仅涉及符号扩散块且无耦合项的表达式以及切线核的舒尔补。由此得出四个结果,即量化筛选耦合能量的精确恒等式、由核的典型相关性控制的训练率退化定律、根据佩克莱数的条件数界以及时间频率在筛选机制中的不参与。条件数的增长减缓了共享步长梯度下降,而连续流不受影响,这使得训练困难归因于优化器而非近似。最后表明,只要架构将与对流和扩散组件相关的参数分开,Adam优化器通过其自适应缩放减轻了这种困难,证实了第二次筛选得出的架构规定。预测在二维交换器上通过数值验证到机器精度,并通过有限宽度网络的完全训练得到验证。

英文摘要

We analyze the training dynamics of variational physics-informed neural networks applied to linear coupled parabolic convection--diffusion--reaction systems, called heterogeneous when only a subset of the components undergoes convective transport. In the neural tangent kernel regime, the gradient flow on the space-time variational residuals reduces to a linear differential system whose operator is a Gram matrix built from the space-time symbol of the system operator and the matrix tangent kernel. The main result is a double screening theorem. Under dominant convection, the Schur complement of this matrix relative to the block of convective components converges to an expression that involves only the diffusive block of the symbol, with no coupling term, together with the Schur complement of the tangent kernel. From this we derive four consequences, namely an exact identity quantifying the screened coupling energy, a degradation law for the training rate governed by the canonical correlations of the kernel, a bound on the condition number in terms of the Péclet number, and the non-participation of temporal frequencies in the screening mechanism. The growth of the condition number slows down shared-step gradient descent, whereas the continuous flow suffers no slowdown, which makes the training difficulty attributable to the optimizer rather than to the approximation. We finally show that the Adam optimizer, through its adaptive scaling, mitigates this difficulty provided that the architecture separates the parameters associated with the convective and diffusive components, confirming the architectural prescription that follows from the second screening. The predictions are validated numerically down to machine precision, on a two-dimensional exchanger, and by the full training of finite-width networks.

补充信息

↑