SpSYRK:分布式稀疏矩阵乘法中仅需一半工作量
SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication
浏览论文内容
中文总结 AI 辅助
本文提出利用对称性的分布式稀疏SYRK方法,通过优化划分与通信减少计算通信开销,在Perlmutter超算32节点上实现2倍加速,可直接替换现有分布式SpGEMM相关应用。
中文摘要 AI 辅助
对称秩-k更新(SYRK)即$\boldsymbol{C} = \boldsymbol{A}\boldsymbol{A}^\top$,用于计算矩阵$\boldsymbol{A}$每对行向量的点积,生成格拉姆矩阵$\boldsymbol{C}$。其稀疏变体支撑机器学习、图分析及基因组学中的相似度搜索,包括在单节点无法容纳的大型数据集上的雅卡尔相似度计算。然而,尽管输入与输出具有对称性,现有分布式稀疏矩阵乘法算法(如Sparse SUMMA)仍将稀疏SYRK视为通用乘法处理,未能充分挖掘其性能潜力。本文提出利用对称性的分布式稀疏SYRK方法:该方法将输出的非对角块划分到进程网格的上三角与下三角部分,且仅计算每个对角块的下三角部分,与最先进的分布式SpGEMM相比,减少了各进程的通信与计算量;第二个变体则重新排序通信,避免显式生成$\boldsymbol{A}^\top$。在Perlmutter超级计算机的32个节点上,当本地乘法占运行时间主导时,该算法相较于优化后的Sparse SUMMA实现了2倍的加速;在通信密集型输入下,优势有所收窄,这一依赖关系可通过算术强度的成本模型预测。本文的变体修正了该问题,在高进程数下始终实现更优的扩展性。该方法可作为任何通过分布式SpGEMM例程计算$\boldsymbol{C} = \boldsymbol{A}\boldsymbol{A}^\top$的应用的直接替换,其三角输出可直接供后续操作使用,从而减少计算与内存占用。
英文摘要
The symmetric rank-$k$ update (SYRK), $C = AA^\top$, computes the dot product of each pair of rows of $A$, producing the Gram matrix $C$. Its sparse variant underpins similarity search in machine learning, graph analytics, and genomics, including Jaccard similarity on datasets too large for a single node. Despite the symmetry in its inputs and outputs, existing distributed sparse matrix multiplication algorithms such as Sparse SUMMA treat sparse SYRK as generic multiplication, computing the full output and materializing the explicit transpose even when the calling application uses only one triangle. Prior distributed $AA^\top$ computations in similarity search and genome assembly inherit this overhead from the underlying SpGEMM. This paper presents SpSYRK and CommSpSYRK, two distributed sparse SYRK algorithms that exploit symmetry. The first, SpSYRK, partitions the off-diagonal blocks of the output between the upper and lower triangular regions of the process grid and computes only the lower-triangular part of each diagonal block, halving per-process computation compared with state-of-the-art distributed SpGEMM. The second, CommSpSYRK, further reorders communication to avoid forming $A^\top$, which reduces per-process communication volume. On 32 nodes of the Perlmutter supercomputer, SpSYRK achieves a 2$\times$ speedup over an optimized Sparse SUMMA on matrices where local multiplication dominates the runtime; the advantage narrows on communication-bound inputs, a dependence that the cost model predicts from the arithmetic intensity. CommSpSYRK fixes this and consistently achieves superior scaling at high process counts. The approach is a drop-in replacement for any application computing $C = AA^\top$ via a distributed SpGEMM routine, and its triangular output can be consumed directly by subsequent operations, reducing both computation and memory footprint.