arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AWE:FP4张量核心上以更少GEMM实现精确整数矩阵乘法的自适应权重编码

AWE: Adaptive Weight Encoding for Exact Integer Matrix Products with Fewer GEMMs on FP4 Tensor Cores

Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri

arXiv 2609.24519首次发表:更新:

发表机构

Graduate School of Informatics Nagoya University; Information Technology Center Nagoya University(名古屋大学信息学研究科; 名古屋大学信息技术中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对FP4张量核心,提出自适应权重编码(AWE),通过自由选择肢权重和线性组合存储平面,将INT8×INT8精确整数矩阵乘法的GEMM次数从9次降至6次,INT4×INT8降至4次,并减少Ozaki方案II的乘积数。

AI 中文摘要

高精度浮点矩阵乘法的模拟,如Ozaki方案,将输入拆分为低精度分量并两两相乘。这些乘积必须无误差,且每个乘积都是整数矩阵乘积乘以一个比例因子。FP4张量核心在NVIDIA B200和B300上速度最快,但无法容纳INT8操作数。乘以2后的FP4值构成集合S = {0, ±1, ±2, ±3, ±4, ±6, ±8, ±12},该集合包含模13的每个余数,因此通过进位,任何整数都可以拆分为FP4可存储的基13数字。先前的工作将每个INT8操作数拆分为3个这样的数字(肢),权重为(1, 13, 169),并两两相乘,对于INT8×INT8需要9次FP4矩阵乘法(GEMM)。经典的减少乘积的方法,如Karatsuba和Toom-Cook方法,不能直接应用:肢的和达到±24,超出集合S。本文探讨一次整数矩阵乘积需要多少次FP4 GEMM。我们提出自适应权重编码(AWE):肢采用自由选择的整数权重,存储的平面是肢的线性组合,精确乘积是FP4 GEMM按重构系数缩放后的总和。对于每个输入范围,我们搜索了这些选择以找到乘积更少的编码,发现INT8×INT8需要6次乘积,INT4×INT8需要4次。该公式也适用于模m,覆盖了Ozaki方案II的余数系统:对于FP64尾数,先前工作的75次乘积减少到59次。有残差和无残差编码在乘积数量上的边界位于输入宽度15附近。我们发布了所找到的编码。

英文摘要

Emulation of high-accuracy floating-point matrix multiplication, as in the Ozaki scheme, splits the inputs into low-precision components and multiplies them pairwise. These products must be error-free, and each is an integer matrix product times a scale factor. FP4 Tensor Cores are the fastest on the NVIDIA B200 and B300 but cannot hold INT8 operands. The FP4 values scaled by 2 form the set $S = \{0, \pm1, \pm2, \pm3, \pm4, \pm6, \pm8, \pm12\}$, which contains every residue modulo 13, so with carries any integer splits into base-13 digits that FP4 can store. Prior work splits each INT8 operand into 3 such digits (limbs) with weights $(1, 13, 169)$ and multiplies them pairwise, 9 FP4 matrix multiplications (GEMMs) for INT8$\times$INT8. The classical ways to reduce products, such as the Karatsuba and Toom--Cook methods, do not apply as they stand: sums of limbs reach $\pm 24$ and leave $S$. This paper asks how many FP4 GEMMs are needed for one integer matrix product. We propose Adaptive Weight Encoding (AWE): the limbs take freely chosen integer weights, the stored planes are linear combinations of limbs, and the exact product is the sum of the FP4 GEMMs scaled by reconstruction coefficients. For each input range, we searched these choices for encodings with fewer products and found INT8$\times$INT8 in 6 products and INT4$\times$INT8 in 4. The formulation also holds modulo $m$, which covers the residue number systems of Ozaki scheme II: for the FP64 significand, the 75 products of prior work are reduced to 59. The boundary in product count between encodings with and without residues lies near input width 15. We release the encodings found.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑