arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20650cs.DCcs.PF

PortLBM:一种利用SYCL在AMD、NVIDIA和英特尔GPU上的便携式格子玻尔兹曼工具

PortLBM: A Portable Lattice Boltzmann Tool Leveraging SYCL on AMD, NVIDIA, and Intel GPUs

Alexander Strack, Marcel Graf, Alexander Van Craen, Dirk Pflüger

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对GPU加速下LBM的跨平台可移植性问题,提出基于SYCL的PortLBM框架,评估其在不同GPU架构上多种数据布局和算法变体对模拟吞吐量的影响,强调在异构计算环境中便携式LBM软件的必要性。

中文摘要 AI 辅助

格子玻尔兹曼方法(LBM)是用于介观尺度模拟流体流动的成熟方法。随着摩尔定律停滞,高性能计算转向GPU加速器,需要确保跨不同硬件平台的可移植性和效率的编程模型。我们提出PortLBM,一个基于SYCL构建的可扩展便携式LBM框架,集成跨平台GPU支持与交互式实时可视化。PortLBM支持从卡门涡街、机翼流到多孔介质等多种模拟场景,易于用新算法和后端扩展。作为性能可移植性研究的一部分,我们在NVIDIA、AMD和英特尔的当代GPU架构上评估PortLBM,考察三种数据布局和四种算法变体对模拟吞吐量的影响。结果表明没有单一配置能在所有GPU供应商上实现最优性能,证实了特定系统调优的必要性。流布局最大化带宽,在当代NVIDIA和英特尔GPU上表现最佳,而束布局提高缓存效率,在AMD GPU上表现出色。双晶格方案实现更高吞吐量,单晶格方案在内存受限下更可取。我们的工作强调了在日益异构的计算环境中适应性强、便携式LBM软件的必要性。

英文摘要

The lattice Boltzmann method (LBM) is a well-established approach for simulating fluid flows at the mesoscopic scale. With the stagnation of Moore's law, high-performance computing has shifted toward GPU accelerators, necessitating programming models that ensure both portability and efficiency across diverse hardware platforms. We present PortLBM, an extensible portable LBM framework built on SYCL that integrates cross-platform GPU support with interactive real-time visualization. PortLBM supports diverse simulation scenarios ranging from Kármán vortex streets and wing flows to porous media, and is designed for easy extension with new algorithms and backends. As part of a performance portability study, we evaluate PortLBM on contemporary GPU architectures from NVIDIA, AMD, and Intel, examining the impact of three data layouts (stream, bundle, and collision) and four algorithmic variants on simulation throughput. Our results show that no single configuration achieves optimal performance across all GPU vendors, confirming the need for system-specific tuning. The stream layout maximizes bandwidth and performs best on the contemporary NVIDIA and Intel GPUs, while the bundle layout improves cache efficiency and excels on the AMD GPU. Two-lattice schemes achieve higher throughput while one-lattice schemes are preferable under memory constraints. Our work underscores the necessity for adaptable, portable LBM software in increasingly heterogeneous computing environments.

补充信息

↑