AI 中文总结
研究针对GPU有损压缩中预测与吞吐量权衡问题,提出FSZ。通过跨块预测状态、每块自适应多阶预测和单通道四路评估三项创新,在单个CUDA内核中提升压缩率,实验显示其压缩率优于同类,且吞吐量最高。
AI 中文摘要
现有的快速GPU误差受限有损压缩器通过纯GPU单核设计实现了高吞吐量,但由于通常在独立块上应用固定的一阶预测器,其压缩率仍然有限。我们提出了FSZ,一种GPU误差受限有损压缩器,通过三项相互强化的算法创新重新设计预测阶段,在单个CUDA内核中实现更高的压缩率和更高的吞吐量:跨块预测状态在256元素块内跨块边界携带洛伦佐预测状态,消除8个边界残差中的7个;每块自适应多阶预测和居中从一阶、二阶和居中变体中为每个块自适应选择最佳压缩策略;单通道四路评估利用有限差分的数学属性从单次数据读取中评估所有变体。在NVIDIA GH200 GPU上对8个实际应用数据集进行的实验表明,FSZ在压缩率方面比cuSZp-P高出10.95倍,比最新的cuSZp-O高出2.92倍。值得注意的是,这些提升并没有带来吞吐量的损失:FSZ在所有评估的压缩器中同时实现了最高的平均吞吐量(压缩676 GB/s,解压缩785 GB/s)。
英文摘要
Existing fast GPU error-bounded lossy compressors have achieved high throughput through pure-GPU single-kernel designs, but their compression ratios remain limited because they typically apply a fixed first-order predictor on independent blocks. We propose FSZ, a GPU error-bounded lossy compressor that redesigns the prediction stage with three mutually reinforcing algorithmic innovations to achieve both higher compression ratios and higher throughput within a single CUDA kernel: (1) cross-block prediction state carries Lorenzo prediction state across block boundaries within 256-element tiles, eliminating 7 out of 8 boundary residuals that inflate encoding rates; (2) per-tile adaptive multi-order prediction and centering adaptively selects the best compression strategy per tile from first-order, second-order, and centering variants; and (3) a single-pass four-way evaluation exploits a mathematical property of finite differences to evaluate all variants from a single data read, enabling richer prediction within the same bandwidth budget as a fixed predictor. Experiments on NVIDIA GH200 GPU with 8 real-world application datasets show that FSZ outperforms cuSZp-P by up to 10.95x and the state-of-the-art cuSZp-O by up to 2.92x in compression ratio. Notably, these gains come with no throughput penalty: FSZ simultaneously achieves the highest average throughput (676 GB/s compression, 785 GB/s decompression) among all evaluated compressors.
CommentsAccepted at SC26