AI 中文总结
提出一种格子玻尔兹曼方法中的梯度重构技术,从平衡态矩结构代数逆推梯度,避免额外输运场,在多种系统中保持二阶精度,并显著提升内存受限场景下的性能。
AI 中文摘要
自动推导将声明的守恒律系统转化为格子玻尔兹曼格式,为每个守恒物理量赋予一组q个种群,其线性平衡态将物理通量嵌入其第一矩中。当通量依赖于守恒状态的梯度时,这些梯度通过将其作为额外的输运场进行追踪来提供。由于格子玻尔兹曼通常受内存限制,这些额外的自由度降低了可实现的吞吐量。为了恢复吞吐量,我们从平衡态的矩结构重构梯度,而不是输运它们。在一阶精度下,守恒量非平衡部分的第一矩携带其梯度。因为重构的通量进入其自身的平衡参考,该矩是梯度在线性算子(由扩散通量雅可比矩阵构建)下的像。重构是该算子的代数逆,适用于一般的梯度形式本构闭包,并从每个声明的通量自动生成。所得格式仅携带守恒量,从种群中重构所需梯度并局部形成通量。在平流-扩散-反应、Allen-Cahn、Navier-Stokes、电阻磁流体动力学和均质化可压缩Navier-Stokes-Fourier系统中,以双精度二阶收敛,与相同分辨率下的梯度追踪精度匹配。在NVIDIA RTX A5000上,单精度下速度提升高达3.7倍,内存受限内核达到峰值内存带宽的97%。
英文摘要
The automatic derivation turns a declared system of conservation laws into a lattice Boltzmann scheme, giving each conserved physical quantity a set of q populations whose linear equilibrium embeds the physical flux in their first moment. When the flux depends on gradients of the conserved state, those gradients are supplied by tracking them as additional transported fields. Since lattice Boltzmann is commonly memory-bound, these additional degrees of freedom reduce the achievable throughput. To reclaim it, we reconstruct the gradients from the moment structure of the equilibrium instead of transporting them. To leading order, the first moment of a conserved quantity's non-equilibrium part carries its gradient. Because the reconstructed flux enters its own equilibrium reference, that moment is the image of the gradient under a linear operator built from the diffusive-flux Jacobian. The reconstruction is that operator's algebraic inverse, generic across gradient-form constitutive closures and generated automatically from each declared flux. The resulting scheme carries the conserved quantities alone, reconstructing the required gradients from the populations and forming the fluxes locally. It converges at second order in double precision across advection-diffusion-reaction, Allen-Cahn, Navier-Stokes, resistive magnetohydrodynamics and homogenized compressible Navier-Stokes-Fourier systems, matching the accuracy of gradient tracking at equal resolution. On an NVIDIA RTX A5000 it is up to 3.7 times faster in single precision, the memory-bound kernels reaching up to 97% of peak memory bandwidth.