LBFAST:面向多GPU架构的轻量级矩表示格子玻尔兹曼求解器
LBFAST: A Lightweight Moment-Represented Lattice Boltzmann Solver for Multi-GPU Architectures
浏览论文内容
中文总结 AI 辅助
LBFAST提出一种基于轻量级矩表示的GPU格子玻尔兹曼求解器,通过减少内存占用实现大规模三维模拟,在多GPU上近理想弱扩展至512个GPU,兼顾精度与能效。
中文摘要 AI 辅助
我们提出了LBFAST,一种基于轻量级矩表示公式的面向GPU的格子玻尔兹曼求解器,其中碰撞后粒子分布函数从一组缩减的矩中即时重建,而非显式存储。该方法显著降低了内存占用,使得在现代加速器架构(其中显存容量和带宽是关键资源)的约束下能够进行大规模三维模拟。该方法通过标准的单组分和双组分基准测试进行了评估,显示出良好的精度和稳定性。在多GPU系统上进行的大量扩展实验表明,在多达512个GPU上实现了近理想的弱扩展性,并在不同速度集上保持了持续的性能。内存使用减少、吞吐量具有竞争力以及能效稳定的组合,使得所提出的公式成为在当前和新兴HPC平台上进行大规模格子玻尔兹曼模拟的实用途径。
英文摘要
We present LBFAST, a GPU-oriented lattice Boltzmann solver based on a lightweight moment-represented formulation, in which post-collision populations are reconstructed on the fly from a reduced set of moments rather than stored explicitly. This approach significantly lowers the memory footprint, enabling large three-dimensional simulations within the constraints of modern accelerator architectures, where VRAM capacity and bandwidth are critical resources. The method is assessed through standard single- and two-component benchmarks demonstrating good accuracy and stability. Extensive scaling experiments on multi-GPU systems show near-ideal weak scaling up to 512 GPUs and sustained performance across different velocity sets. The combination of reduced memory usage, competitive throughput, and stable energy efficiency makes the proposed formulation a practical route for large-scale lattice Boltzmann simulations on current and emerging HPC platforms.
发表机构
- Istituto per le Applicazioni del Calcolo, Consiglio Nazionale delle Ricerche(意大利国家研究委员会计算应用研究所)
- Roma Tre University(罗马第三大学)
- CINECA
- NVIDIA Development UK Ltd(英伟达英国开发有限公司)
- Center for Life Nano- & Neuro-Science, Fondazione Istituto Italiano di Tecnologia(意大利理工大学基金会生命纳米与神经科学中心)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。