Ozaki Scheme II 在 CPU 上同样快速:Intel AMX-INT8 与 Arm SVE2-i8mm 上的多精度矩阵乘法
Ozaki Scheme II Is Fast on CPUs Too: Multiple-Precision Matrix Multiplication on Intel AMX-INT8 and Arm SVE2-i8mm
浏览论文内容
中文总结 AI 辅助
本文实现并优化了 Ozaki Scheme II 多精度矩阵乘法,在 Intel AMX-INT8 和 Arm SVE2-i8mm 上分别构建后端,实现高达 167 倍加速且精度在 1 ulp 以内,并定量解释了其在 CPU 上高效的原因。
中文摘要 AI 辅助
我们实现了 Ozaki Scheme II(残数系统 + 中国剩余定理),该方案将多精度稠密矩阵乘法简化为 CPU 上的一系列低精度、高吞吐量的整数或浮点 GEMM 运算。两个后端构建在共享的 CRT 重构阶段之上:(a) 在 Intel AMX 上使用精确的 INT8 x INT8 -> INT32 分块乘积,以及 (b) 二进制64 DGEMM,即原始 Ozaki Scheme II 论文的 CPU 实现。在双插槽 Xeon Gold 6526Y(Emerald Rapids,32 核)上,我们评估了 53-2048 位的有效数精度和矩阵维度 N = 256-8192。结果始终在高精度 MPFR 参考值的 1 ulp 以内(本质上为正确舍入),同时运行速度比朴素 MPFR 矩阵乘积快高达 167 倍,比 BNCmatmul 的 Strassen 乘法快高达 588 倍,比 Ozaki Scheme I(FP64 切片 + OpenBLAS DGEMM)快 9-78 倍。两个后端之间的盈亏平衡点大约在 N = 2048:低于该值时,二进制64 后端因其模数较少而胜出;高于该值时,AMX-INT8 后端因 GEMM 占主导而胜出。我们进一步将实现移植到 AArch64(NVIDIA GB10:Cortex-X925 x 10 + Cortex-A725 x 10)。由于该机器缺乏 SME/SME2,INT8 内核使用 SVE2 i8mm 扩展的 SMMLA 矩阵乘积指令。我们获得了持续 6.5 TOPS 的精确 INT8 GEMM,并且在所有条件下仍保持在 1 ulp 以内,相对于 BNCmatmul 的 Ozaki Scheme I(OpenBLAS 链接例程)加速比为 14-89 倍,相对于公平调整的 OzI-best 变体加速比为 6-19 倍。论文还包括对 Ozaki Scheme II 的教程式介绍(“Introduction to Ozaki Scheme II” 一节),并定量解释了为什么这种看似 GPU 专用的技术在 CPU 上同样快速。
英文摘要
We implement Ozaki Scheme II (residue number system + Chinese remainder theorem), which reduces multiple-precision dense matrix multiplication to a sequence of low-precision, high-throughput integer or floating-point GEMMs on CPUs. Two backends are built on top of a shared CRT reconstruction stage: (a) exact INT8 x INT8 -> INT32 tile products on Intel AMX, and (b) binary64 DGEMM, the CPU construction of the original Ozaki Scheme II paper. On a two-socket Xeon Gold 6526Y (Emerald Rapids, 32 cores), we evaluate significand precisions of 53-2048 bits and matrix dimensions N = 256-8192. The results are always within 1 ulp of a high-precision MPFR reference (essentially correctly rounded), while running up to 167x faster than a naive MPFR matrix product, up to 588x faster than BNCmatmul's Strassen multiplication, and 9-78x faster than Ozaki Scheme I (FP64 slicing + OpenBLAS DGEMM). The break-even point between the two backends is approximately N = 2048: below it, the binary64 backend wins thanks to its smaller number of moduli; above it, the AMX-INT8 backend wins as the GEMMs dominate. We further port the implementation to AArch64 (NVIDIA GB10: Cortex-X925 x 10 + Cortex-A725 x 10). Since this machine lacks SME/SME2, the INT8 kernel uses the SMMLA matrix-product instruction of the SVE2 i8mm extension. We obtain an exact INT8 GEMM sustaining 6.5 TOPS and, still within 1 ulp across all conditions, speedups of 14-89x over BNCmatmul's Ozaki Scheme I (OpenBLAS-linked routine) and 6-19x over a fairness-adjusted OzI-best variant. The paper also includes a tutorial introduction to Ozaki Scheme II (Section "Introduction to Ozaki Scheme II") and a quantitative explanation of why this seemingly GPU-specific technique is fast on CPUs as well.
发表机构
- Otemon Gakuin University(大东文化大学)
机构由 AI 辅助整理,请以论文原文为准。