arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$g$MAGNUS:通过分层多分裂在GPU上对不规则矩阵进行快速稀疏矩阵乘法

$g$MAGNUS: Fast SpGEMM on GPUs for Irregular Matrices via Hierarchical Multisplit

Jordi Wolfson-Pou, Ahmed Helal, Fabrizio Petrini

arXiv 2607.22866首次发表:更新:

AI 中文总结

研究针对GPU上不规则矩阵的稀疏矩阵乘法问题,提出$g$MAGNUS算法,通过行内重新排序等操作,自动确定参数,实验表明其比五种领先算法有显著加速,核心内核接近理论峰值性能。

AI 中文摘要

我们提出了$g$MAGNUS,一种用于在GPU上对不规则矩阵进行稀疏矩阵乘法(SpGEMM)的新算法。此类矩阵通常包含许多“重行”,即中间积大到迫使本地内存累加器溢出到全局内存的行。$g$MAGNUS通过计算中间积的行内重新排序来解决此问题,将重行细分为可在本地内存中完全累加的独立块。这种重新排序使用了新颖的外积和分层多分裂操作。该算法具有输入和系统感知能力,根据输入矩阵维度和本地内存大小自动确定块数和多分裂级别。在两个广泛数据集上的实验结果表明,$g$MAGNUS在英特尔Ponte Vecchio和英伟达H200上比五种领先算法(包括MKL和cuSPARSE)实现了1.81至7.62倍的几何平均加速。此外,对$g$MAGNUS的核心内核进行了评估,与理论上限相比实现了接近峰值的性能。

英文摘要

We present $g$MAGNUS, a novel algorithm for sparse matrix-matrix multiplication (SpGEMM) of irregular matrices on GPUs. Such matrices often contain many heavy rows, those with large intermediate products that force local memory accumulators to spill to global memory. $g$MAGNUS addresses this by computing an intra-row reordering of intermediate products, subdividing heavy rows into independent chunks that can be accumulated completely in local memory. This reordering uses novel outer product and hierarchical multisplit operations. The algorithm is input- and system-aware, automatically determining the number of chunks and multisplit levels based on the input matrix dimensions and local memory size. Experimental results on two extensive datasets show that $g$MAGNUS achieves a geometric-mean speedup of 1.81 to 7.62$\times$ over five leading algorithms (including MKL and cuSPARSE) on Intel Ponte Vecchio and NVIDIA H200. Additionally, the core kernels of $g$MAGNUS are evaluated, achieving near-peak performance compared to their theoretical upper bound.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑