AI 中文总结
提出草图切线一致性损失(sTCL),通过随机草图扰动实现PDE神经算子的在线导数信息训练,无需离线标签,在多个PDE上达到与离线方法相当的精度。
AI 中文摘要
导数信息训练通过直接监督神经算子的输入-输出敏感性来改进神经算子,这在神经算子被用作反问题、PDE约束优化、设计和控制的可微代理时至关重要。然而,现有方法依赖离线生成的导数标签,导致数据生成缓慢、存储密集,且难以跨数据集、分辨率或扰动基进行适应。我们提出了草图切线一致性损失(sTCL),一种在线导数信息训练目标,对于具有已知且可微残差的PDE,它直接从控制方程强制敏感性一致性,无需离线切线标签或神经算子架构更改。sTCL在训练期间使用随机草图的输入扰动来提供轻量级的导数级物理约束。然而,原始的向前敏感性残差惩罚可能对刚性、病态、不定或耦合鞍点切线算子失效。为解决此问题,我们引入了由简单切线算子决策规则选择的轻量级算子感知损失调节机制。在Helmholtz、非线性扩散-反应、Burgers、Allen-Cahn和Navier-Stokes方程上,所有方法使用相同的神经算子骨干,特定于PDE的sTCL损失实现了与离线导数信息训练(DIFNO)相当的解和雅可比精度,同时消除了离线导数数据生成和存储流程。这些结果表明,在线导数信息训练不必仅仅将离线切线求解成本摊销到训练中;通过适当的草图和损失设计,sTCL提供了一条有效的即插即用路径来实现导数信息神经算子。代码可在以下网址获取:https URL。
英文摘要
Derivative-informed training improves neural operators by directly supervising their input-output sensitivities, which is crucial when neural operators are used as differentiable surrogates for inverse problems, PDE-constrained optimization, design, and control. However, existing methods rely on offline-generated derivative labels, making data generation slow, storage-intensive, and difficult to adapt across datasets, resolutions, or perturbation bases. We propose sketched tangent consistency loss (sTCL), an on-the-fly derivative-informed training objective that, for PDEs with a known and differentiable residual, enforces sensitivity consistency directly from the governing equation without offline tangent labels or neural-operator architecture changes. sTCL uses randomly sketched input perturbations to provide a lightweight derivative-level physics constraint during training. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, indefinite, or coupled saddle-point tangent operators. To address this, we introduce lightweight operator-aware loss-conditioning mechanisms selected by a simple tangent-operator decision rule. Across Helmholtz, nonlinear diffusion-reaction, Burgers, Allen-Cahn, and Navier-Stokes, with the same neural-operator backbone for all methods, the PDE-specific sTCL losses achieve solution and Jacobian accuracy comparable to offline derivative-informed training (DIFNO) while eliminating the offline derivative-data generation and storage pipeline. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and loss design, sTCL provides an effective drop-in path to derivative-informed neural operators. Code is available at https://github.com/yang9579/Derivative-informed-traning-on-the-fly.
Comments21 pages, 1 figure, 9 tables