arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于FP4张量核心的Ozaki格式I/II的DGEMM:基-13 E2M1 limb表示

DGEMM with Ozaki Scheme I/II on FP4 Tensor Cores: A Base-13 E2M1 Limb Representation

Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri

arXiv 2608.06812首次发表:更新:

AI 中文总结

该研究在FP4张量核心上构建Ozaki格式I/II,提出基-13 E2M1 limb表示方法及内核优化,在RTX PRO 6000 Blackwell上实现FP64矩阵乘法模拟,大问题规模下性能超FP8版本。

AI 中文摘要

本文提出一种方法及其实现,用于在FP4(E2M1:2位指数位和1位尾数位)张量核心上构建Ozaki格式I和II,以在低精度算术单元上实现高精度矩阵乘法,从而模拟FP64矩阵乘法(DGEMM)。此前的实现基于INT8和FP8,而使用速度更快的FP4尚未实现。其关键特性在于,每个FP4值在加倍后会变为整数,且将该整数集按13的倍数偏移可覆盖所有整数。利用该特性将任意整数转换为基-13 FP4 limb,可在FP32累加器中保持中间和无误差,这使得FP4张量核心可用于Ozaki格式I和II。基于相同原理,INT8张量核心的整数GEMM也可在FP4张量核心上实现位精确模拟。当FP4张量核心的吞吐量是FP8的两倍时,FP4上的Ozaki格式II理论上实现的性能略高于其FP8对应版本。本文进一步提出内核实现优化,以提高峰值性能的达到比例,在理论优势之上获得实测加速。我们在RTX PRO 6000 Blackwell上验证了这一点,实现的性能与现有基于FP8的Ozaki格式II实现相当,且在大问题规模(16384³)下实际超过了该实现。

英文摘要

This paper proposes a method and its implementation for emulating FP64 matrix multiplication (DGEMM) by constructing, on FP4 (E2M1; 2 exponent bits and 1 mantissa bit) Tensor Cores, Ozaki schemes I and II, which realize high-precision matrix multiplication on low-precision arithmetic units. Prior implementations were based on INT8 and FP8, and the use of the faster FP4 had not been realized. The key property is that every FP4 value becomes an integer when doubled, and that shifting this integer set by multiples of 13 covers all integers. Converting an arbitrary integer into base-13 FP4 limbs by this property keeps intermediate sums error-free in FP32 accumulators, which makes FP4 Tensor Cores usable for Ozaki schemes I and II. By the same principle, the integer GEMM of INT8 Tensor Cores can also be emulated bit-exactly on FP4 Tensor Cores. When FP4 Tensor Cores have twice the throughput of FP8, Ozaki scheme II on FP4 theoretically achieves slightly higher performance than its FP8 counterpart. This paper further proposes kernel implementation optimizations raising the attained fraction of peak performance, obtaining a measured speedup on top of the theoretical advantage. We verify this on an RTX PRO 6000 Blackwell, achieving performance competitive with that of an existing FP8-based implementation of Ozaki scheme II, and actually exceeding it at a large problem size ($16384^3$).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑