arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在超算系统上加速相场模拟,实现快于实时的沉淀析出老化预测

Accelerating phase-field simulations on exascale computing systems for faster-than-real-time precipitate aging predictions

Stephen DeWitt, David J. Gardner, Philip Fackler, Yonggil Song, Miroslav Stoyanov, Carol S. Woodward, Balasubramaniam Radhakrishnan

arXiv 2609.36100首次发表:更新:

发表机构

Oak Ridge National Laboratory; Lawrence Livermore National Laboratory(橡树岭国家实验室; 劳伦斯利弗莫尔国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种结合GPU计算、分布式并行和高阶时间积分的性能可移植方法,在MEUMAPPS框架中加速傅里叶伪谱相场模拟,实现Ni-Nb-Fe合金沉淀析出老化的快于实时预测,并显著提升计算速度与可扩展性。

AI 中文摘要

三维相场模拟是材料微观结构预测的金标准,但其计算成本往往将应用相关计算限制在适中的域大小和时间尺度。在此,我们提出了一种整体方法,通过结合性能可移植的GPU计算、大规模分布式内存并行性以及MEUMAPPS C++框架内的高阶隐式-显式时间积分,加速傅里叶伪谱相场模拟。我们通过一个15亿网格点的模拟来演示该方法,模拟了Ni-Nb-Fe合金中1,920个γ''沉淀相的生长和粗化。模拟的七小时热处理在5.5小时内完成,使得计算快于实时,并且估计比仅使用CPU的一阶基线快217-501倍。这一性能使得数千个相互作用的沉淀相的三维模拟变得可行,从而能够对集体微观结构现象进行定量研究,构建大规模模拟集成,并实时集成到控制和解释实验中。基准测试显示,相对于相当的CPU资源,单节点GPU加速高达17.5倍,对于120亿网格点的问题,在4,096个GPU上实现近理想的强扩展,并且在科学相关的误差容限下,通过四阶时间积分进一步加速2.6-2.9倍。这些结果建立了一种性能可移植的策略,用于利用领先规模的GPU系统,该策略适用于超越相场模型的广泛傅里叶伪谱模拟,如流体动力学和晶体塑性模拟。

英文摘要

Three-dimensional phase-field simulations are a gold-standard for microstructure prediction for materials, but their computational cost often limits application-relevant calculations to modest domain sizes and timescales. Here, we present a holistic approach for accelerating Fourier pseudospectral phase-field simulations by combining performance-portable GPU computing, large-scale distributed-memory parallelism, and high-order implicit-explicit time integration within the MEUMAPPS C++ framework. We demonstrate this approach with a 1.5-billion-grid-point simulation of the growth and coarsening of 1,920 gamma'' precipitates in a Ni-Nb-Fe alloy. The simulated seven-hour heat treatment is completed in 5.5 hours, making the calculation faster than real time and an estimated 217-501x faster than a CPU-only, first-order baseline. This performance makes three-dimensional simulations of thousands of interacting precipitates tractable, enabling quantitative studies ocollective microstructural phenomena, large simulation ensembles, and real-time integration into controlling and interpreting experiments. Benchmarking shows single-node GPU speedups of up to 17.5x relative to comparable CPU resources, near-ideal strong scaling to 4096 GPUs for a 12-billion-grid-point problem, and a further 2.6-2.9x acceleration from fourth-order time integration at scientifically relevant error tolerances. These results establish a performance-portable strategy for exploiting leadership-scale GPU systems that is applicable to a broad class of Fourier pseudospectral simulations beyond phase-field models such as fluid dynamics and crystal plasticity simulations.

Comments39 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑