arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Diff-NekRS:一种用于多时间步求解器在环训练的可扩展可微框架

Diff-NekRS: A Scalable Differentiable Framework for Multi-Timestep Solver-in-the-Loop Training

Junoh Jung, Riccardo Balin, Bethany Lusch, Emil Constantinescu

arXiv 2609.23208首次发表:更新:

发表机构

Argonne National Laboratory(阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Diff-NekRS可微框架,将神经修正嵌入NekRS求解器,实现多时间步求解器在环训练,在保持速度优势的同时提升粗网格轨迹精度。

AI 中文摘要

混合物理-机器学习求解器通过将可训练修正嵌入时间积分来改进欠分辨模拟。在自回归推理过程中,重复的求解器-模型交互会放大微小误差,这促使了多时间步求解器在环训练的需求。然而,生产级求解器很少暴露用于通过此类展开进行反向传播所需的导数。我们提出了Diff-NekRS,一种可扩展的可微框架,它将神经修正直接嵌入到GPU加速的NekRS不可压缩流求解器中。NekRS计算权威的正向轨迹,手动实现的精确离散伴随对支持的完全离散时间步进行微分,而LibTorch提供神经向量-雅可比乘积和参数梯度。端到端的泰勒和中心有限差分测试验证了二维圆柱绕流(2Dcyl)和三维泰勒-格林涡(3DTGV)在五个时间跨度和12-1020个MPI秩上的组装梯度。在1020个秩下,启用优化器的设置后训练更新分别保留了2Dcyl和3DTGV的54.5%-78.0%和80.7%-81.9%的弱扩展效率。在200步自回归推理中,相对于未修正的粗网格P=2基线,M=50模型将2Dcyl和3DTGV的三种子中值终端相对L2速度误差分别降低了59.2%和12.1%,并且相对于相应的P=7配置,在相同的模拟时间间隔内分别保持了5.38倍和2.49倍的墙钟加速。这些结果建立了一条经过验证且可扩展的多时间步求解器在环训练路径,该路径提高了粗网格轨迹精度,同时保持相对于高阶参考的速度优势。

英文摘要

Hybrid physics-machine-learning solvers improve under-resolved simulations by embedding trainable corrections into the time integration. During autoregressive inference, repeated solver-model interactions can amplify small errors, motivating multi-timestep solver-in-the-loop training. However, production solvers rarely expose the derivatives needed to backpropagate through such rollouts. We introduce Diff-NekRS, a scalable differentiable framework that embeds neural corrections directly in the GPU-accelerated NekRS incompressible-flow solver. NekRS computes the authoritative forward trajectory, a manually implemented exact discrete adjoint differentiates the supported fully discrete timestep, and LibTorch supplies neural vector-Jacobian products and parameter gradients. End-to-end Taylor and centered finite-difference tests verify the assembled gradient for two-dimensional cylinder flow (2Dcyl) and the three-dimensional Taylor-Green vortex (3DTGV) across five horizons and 12-1,020 MPI ranks. At 1,020 ranks, optimizer-enabled post-setup training updates retain 54.5%-78.0% and 80.7%-81.9% weak-scaling efficiency for 2Dcyl and 3DTGV, respectively. In 200-step autoregressive inference, the M = 50 model reduces the three-seed median terminal relative L2 velocity error by 59.2% for 2Dcyl and 12.1% for 3DTGV relative to the uncorrected coarse-grid P = 2 baseline, and retains wall-clock speedups of 5.38x and 2.49x, respectively, relative to the corresponding P = 7 configurations for equal simulated-time intervals. These results establish a verified and scalable path for multi-timestep solver-in-the-loop training that improves coarse-grid trajectory accuracy while retaining a speed advantage over the high-order reference

Comments22 pages, 8 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑