arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FP64 是你想要的,INT8 是你需要的,FP4/6/8 是你拥有的

FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have

Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler

arXiv 2609.37693首次发表:更新:

发表机构

ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种组合优化方法,为 Ozaki 系列低精度矩阵乘法方案选择最优模数组合,最小化 GEMM 数量,并在 Blackwell GPU 上实现 FP4/6/8 方案,比原生 FP64 快最多 83 倍。

AI 中文摘要

Ozaki 方案 II 通过模两两互素的余数,用 INT8 矩阵乘积模拟 FP64 矩阵乘积,随后也出现了针对 FP8 和 FP4 的变体。我们将这些方案视为一个家族,并将方案的选择视为一个组合优化问题,该问题最小化低精度 GEMM(通用矩阵乘法)的数量。给定每个模数下,由低精度 GEMM 计算模乘积的有限种方式,我们针对任何格式、累加器和内部维度,找到使用最少 GEMM 的模数及计算方式的选择,并推导出整个家族中 GEMM 数量的下界。将该方法应用于当前 GPU 的格式,得到了首个 FP6 方案、一个比以往任何方案使用更少 GEMM 的 FP8 方案,以及一个 FP4 方案,下界表明该方案在模数位于指定范围内的任何家族方案中所需的 GEMM 数量最少。在三个 Blackwell GPU 上实现后,INT8、FP8 和 FP4 方案比原生 FP64 运行得更快,在 B300 上最高加速达 83 倍。

英文摘要

Ozaki scheme II emulates FP64 matrix products with INT8 ones through residues modulo pairwise coprime moduli, and variants for FP8 and FP4 have followed. We treat these schemes as one family and pose the choice of a scheme as a combinatorial program that minimizes the number of low-precision GEMMs. Given, for each modulus, a finite set of ways to compute products modulo it from low-precision GEMMs, we find the choice of moduli and ways with the fewest GEMMs, for any format, accumulator and inner dimension, and derive lower bounds on the GEMM count over the whole family. Applied to the formats of current GPUs, the method gives the first FP6 schemes, an FP8 scheme with fewer GEMMs than any previous one, and an FP4 scheme that the bounds show needs the fewest GEMMs of any scheme in the family whose moduli lie in a stated range. Implemented on three Blackwell GPUs, the INT8, FP8 and FP4 schemes run faster than native FP64, up to 83x on B300.

Comments12 pages. Preliminary version; an extended version with the proofs, the search details and further measurements will follow

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑