迈向GPU驻留气候模型:球谐变换通信瓶颈的有损压缩可行性研究
Toward GPU-Resident Climate Models: A Feasibility Study on Lossy Compression for the Spherical Harmonic Transform's Communication Bottleneck
- University of Trento(特伦托大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对气候模型中球谐变换的通信瓶颈,提出GPU驻留有损压缩方案,证明ZFP压缩在相近或更低误差下显著减少通信时间,优于float16截断。
AI中文摘要:
运行中的伪谱大气模型,如ECMWF综合预报系统(IFS),目前几乎完全在CPU上运行;GPU移植正在积极开发中,但尚未用于生产。这些模型依赖球谐变换(SHT)。每个时间步需要前向和逆向SHT,且两次传递都依赖于全局铅笔转置,该转置将多维数组跨计算节点重新分配。在大型节点数下,这些全局集合通信主导墙钟时间。我们研究GPU驻留有损压缩,使用DYAMOND高分辨率运行数据集中的代表性场作为输入,并结合实测GPU压缩吞吐量与SimGrid网络模拟,我们表明,在每值16位(rate-16)的ZFP压缩(约与float16相同的存储预算)下,匹配float16截断的通信时间减少,同时提供约1600倍更低的平均相对误差。在每值8位(rate-8)的ZFP压缩下,实现约1.93倍的float16加速比,同时保持4倍更低的平均相对误差。
英文摘要:
Operational pseudospectral atmospheric models such as the ECMWF Integrated Forecasting System (IFS) run today almost exclusively on CPUs; GPU ports are under active development but not yet used in production. These models rely on the Spherical Harmonic Transform (SHT). Each time-step requires forward and inverse SHTs, and both passes depend on global pencil transposition that redistribute multi-dimensional arrays across compute nodes. At large node counts these global collectives dominate wall-clock time. We investigate GPU-resident lossy compression, using representative fields from the DYAMOND high-resolution operational dataset as input, and combining measured GPU compression throughput with SimGrid network simulation, we show that ZFP at 16 bits per value (rate-16), about the same storage budget as float16, matches the communication-time reduction of float16 truncation while delivering approximately $\mathbf{1600\times}$ lower mean relative error. ZFP at 8 bits per value (rate-8) achieves approximately $\mathbf{1.93\times}$ the speedup of float16 while retaining $\mathbf{4\times}$ lower mean relative error.