发表机构
RIKEN Center for Computational Science (R-CCS)(理化学研究所计算科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文工程化FP8 Ozaki II解构路径,通过解构感知模型、一次转换工作区及流侧余数转换模式,将Rubin GPU上FP64仿真矩阵乘法下限从235提升至473 TFLOPS,性能翻倍。
AI 中文摘要
FP8 Ozaki II 通过 CRT 余数系统上的张量核心乘积来仿真 FP64 矩阵乘法;在张量指令发出之前,将操作数转换为余数平面(即配套论文《FP8 就是你所需要的一切,第 1 部分》中张量-内存平衡模型的解构项)会消耗整数流水线和内存资源。本文对该路径进行了工程化;所有结果均为模型预测,有待测量验证。首先,提出一个解构感知模型:在 NVIDIA Rubin GPU 上,仿真速率仅在单个线程块簇内达到算术上限 $P_{\ m FP8}/(3r+1)$(在 $r=12$ 时约为 473 TFLOPS);更大的输出会在运行中重新拆分,并保持在约 235 TFLOPS 的下限(即上限的一半,这是三个设计整数的比值,而非拟合结果),而实际求解器的高瘦形状则接近交叉点,目前比简单解构快 1.6-1.9 倍。其次,方法上:采用一次转换的余数工作区、在整数张量流水线(或纯 SIMT dp4a)上的精确双肢常数归约 GEMM,以及将转换与 MMA 流水线化,使交叉点从约 1211 降至约 480-730。第三,模数协同设计:全字节和混合集合、两个供应界限以及尾部模数的进位校正 E4M3 拆分。第四,也是核心,闭式下限指明了其硬件出路,而回报属于 Rubin:在异步复制路径上的一种流侧余数转换模式(选项 C),一个按物料清单计量的窄固定功能块,将平面形成从算术流水线中移出,并将下限从 235 TFLOPS 提升至完整的 473 TFLOPS 上限,同时保持簇覆盖范围不变,使每个 Rubin GPU 的 HPL 级 FP64 性能大约翻倍,并解除转换受限的稀疏内核的束缚。NVIDIA GB300 GPU 的 135 TFLOPS 上限位于其自身下限处,收益甚微;下限及其补救措施均为 Rubin 规模。应用轨迹为分析提供了依据;常量已通过脚本检查。
英文摘要
FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibrium model of the companion paper "FP8 is All You Need, Part 1") costs integer-pipe and memory resources before tensor instructions issue. This paper engineers that path; every result is a model projection pending measurement. First, a deconstruction-aware model: on the NVIDIA Rubin GPU the emulated rate reaches the arithmetic roof $P_{\rm FP8}/(3r+1)$ ($\approx 473$ TFLOPS at $r=12$) only within one thread-block cluster; larger outputs are re-split on the fly and held at a floor of $\approx 235$ TFLOPS (half the roof, a ratio of three design integers, not a fit), while real solvers' tall/skinny shapes stay near the crossover, $1.6$-$1.9\times$ over simple deconstruction today. Second, the method: convert-once residue workspaces, an exact two-limb constant-reduction GEMM on integer tensor pipes (or pure-SIMT dp4a), and conversion pipelined behind the MMAs, moving the crossover from $\approx 1211$ to $\approx 480$-$730$. Third, modulus co-design: all-byte and hybrid sets, two supply bounds and a carry-corrected E4M3 split of tail moduli. Fourth and central, the closed-form floor names its hardware escape, and the prize is Rubin's: a stream-side residue-conversion mode on the asynchronous copy path (Option C), a narrow fixed-function block sized as a bill of materials, takes plane formation off the arithmetic pipes and lifts the floor from 235 TFLOPS to the full 473-TFLOPS roof at unchanged cluster reach, about doubling HPL-class FP64 per Rubin GPU, and unbinds conversion-bound sparse kernels. The NVIDIA GB300 GPU, whose 135-TFLOPS roof sits at its own floor, gains little; floor and remedy are Rubin-scale. Application traces ground the analysis; constants are script-checked.
CommentsEssentially a part 3 paper of the FP8 is all you need work but also standalone work to significantly enhance the Ozaki II scheme as well as hardware assists to further accelerate FP64 emulation