发表机构
Maginfra Co., Ltd.(麦格菲拉有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对CPU提出基于Intel AMX和BF16矩阵乘积的方法,仿真FP32和FP64 GEMM,通过调整分量乘积数实现精度-性能权衡,AMX-FP32吞吐量优于oneMKL SGEMM,AMX-FP64在大阶数下低乘积计数变体性能超DGEMM。
AI 中文摘要
现代CPU日益集成针对低精度AI工作负载优化的高吞吐量矩阵引擎,而许多科学计算应用仍依赖FP32和FP64 GEMM来满足其数值精度要求。这种不匹配催生了一种算法桥梁,即利用低精度矩阵乘积来仿真更高精度的GEMM。本文提出一种面向CPU的方法,基于Intel高级矩阵扩展(Advanced Matrix Extensions, AMX)和BF16矩阵乘积。对于FP32,每个操作数被分解为三个BF16分量,并评估六个选定的分量乘积,相对于oneMKL SGEMM达到FP32级精度,但不追求元素级或位级一致性。对于在支持的BF16指数范围内的FP64输入,该方法采用简化的固定六切片Ozaki分解。每个保留的BF16乘积首先以FP32生成,然后扩展并以FP64累加。四种乘积计数设置分别保留6、10、15或21个分量乘积,相对于oneMKL DGEMM展现出精度-性能权衡。该实现结合了预计算的打包分量缓冲区、VNNI打包的B面板,以及FP32 tile驻留的操作数复用调度。在测试的方阵上,AMX-FP32的吞吐量超过oneMKL SGEMM;对于AMX-FP64,低乘积计数变体在足够大的阶数下可超过DGEMM,而保留更多乘积则以额外成本提升精度。
英文摘要
Modern CPUs increasingly integrate high-throughput matrix engines optimized for low-precision AI workloads, while many scientific computing applications still rely on FP32 and FP64 GEMM to meet their numerical accuracy requirements. This mismatch motivates an algorithmic bridge that uses low-precision matrix products to emulate higher-precision GEMM. This paper presents a CPU-oriented method based on Intel Advanced Matrix Extensions (AMX) and BF16 matrix products. For FP32, each operand is decomposed into three BF16 components and six selected component products are evaluated, targeting FP32-level accuracy relative to oneMKL SGEMM without claiming elementwise or bitwise identity. For FP64 inputs within the supported BF16 exponent range, the method uses a simplified fixed six-slice Ozaki decomposition. Each retained BF16 product is first produced in FP32, then widened and accumulated in FP64. Four product-count settings retain 6, 10, 15, or 21 component products, exposing the accuracy--performance tradeoff relative to oneMKL DGEMM. The implementation combines precomputed packed component buffers, VNNI-packed $B$ panels, and an FP32 tile-resident operand-reuse schedule. On the tested square matrices, AMX-FP32 exceeds oneMKL SGEMM throughput. For AMX-FP64, low-product-count variants can exceed DGEMM at sufficiently large orders, while retaining more products improves accuracy at additional cost.