发表机构
NVIDIA Corporation(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出复数矩阵的2M乘法算法,将整数复数GEMM降至2个实数GEMM,浮点场景结合Ozaki-II方案,性能约为实数GEMM的两倍,并推导出新的SYRK/HERK算法。
AI 中文摘要
复数矩阵乘法通常使用4个同尺寸的实数矩阵乘法(GEMM)来计算。众所周知的3M乘法算法将此成本降低到3个实数GEMM,并伴随二次时间的前处理和后处理步骤。在本文中,我们针对实部和虚部均为整数的矩阵,将3M算法减少为2M算法,即仅使用2个同尺寸的实数GEMM以及二次时间的前处理和后处理步骤来执行复数GEMM。对于浮点矩阵,2M乘法与Ozaki-II方案自然结合,产生一种实用且高性能的算法,其计算复数浮点GEMM的时间大约是同尺寸实数GEMM的两倍。作为推论,我们推导出新的对称秩-k更新(SYRK/HERK)算法,这些算法内部使用完整的矩形GEMM。
英文摘要
Complex matrix multiplication is typically computed using 4 real matrix multiplications (GEMMs) of the same size. The well-known 3M multiplication algorithm reduces this cost to 3 real GEMMs, together with quadratic time pre- and post-processing steps. In this paper, we reduce 3M to 2M for matrices with integer real and imaginary parts, performing complex GEMM with only 2 real GEMMs of the same size, along with quadratic time pre- and post-processing. For floating-point matrices, 2M multiplication combines naturally with the Ozaki-II scheme, yielding a practical, high-performance algorithm for computing a complex floating-point GEMM in roughly twice the time of a real GEMM of the same size. As corollaries, we derive new algorithms for symmetric rank-$k$ updates (SYRK/HERK) that internally use full rectangular GEMMs.
Comments18 pages, 5 figures