arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Ozaki 2.5:在FP8张量核心上工程化fp64仿真稠密矩阵乘法的解构路径

Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores

Satoshi Matsuoka

arXiv 2609.09095首次发表:更新:

发表机构

RIKEN Center for Computational Science (R-CCS)(理化学研究所计算科学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文工程化FP8 Ozaki II解构路径,通过解构感知模型、一次转换工作区及流侧余数转换模式,将Rubin GPU上FP64仿真矩阵乘法下限从235提升至473 TFLOPS,性能翻倍。

AI 中文摘要

FP8 Ozaki II 通过 CRT 余数系统上的张量核心乘积来仿真 FP64 矩阵乘法;在张量指令发出之前,将操作数转换为余数平面(即配套论文《FP8 就是你所需要的一切,第 1 部分》中张量-内存平衡模型的解构项)会消耗整数流水线和内存资源。本文对该路径进行了工程化;所有结果均为模型预测,有待测量验证。首先,提出一个解构感知模型:在 NVIDIA Rubin GPU 上,仿真速率仅在单个线程块簇内达到算术上限 $P_{\ m FP8}/(3r+1)$(在 $r=12$ 时约为 473 TFLOPS);更大的输出会在运行中重新拆分,并保持在约 235 TFLOPS 的下限(即上限的一半,这是三个设计整数的比值,而非拟合结果),而实际求解器的高瘦形状则接近交叉点,目前比简单解构快 1.6-1.9 倍。其次,方法上:采用一次转换的余数工作区、在整数张量流水线(或纯 SIMT dp4a)上的精确双肢常数归约 GEMM,以及将转换与 MMA 流水线化,使交叉点从约 1211 降至约 480-730。第三,模数协同设计:全字节和混合集合、两个供应界限以及尾部模数的进位校正 E4M3 拆分。第四,也是核心,闭式下限指明了其硬件出路,而回报属于 Rubin:在异步复制路径上的一种流侧余数转换模式(选项 C),一个按物料清单计量的窄固定功能块,将平面形成从算术流水线中移出,并将下限从 235 TFLOPS 提升至完整的 473 TFLOPS 上限,同时保持簇覆盖范围不变,使每个 Rubin GPU 的 HPL 级 FP64 性能大约翻倍,并解除转换受限的稀疏内核的束缚。NVIDIA GB300 GPU 的 135 TFLOPS 上限位于其自身下限处,收益甚微;下限及其补救措施均为 Rubin 规模。应用轨迹为分析提供了依据;常量已通过脚本检查。

英文摘要

FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibrium model of the companion paper "FP8 is All You Need, Part 1") costs integer-pipe and memory resources before tensor instructions issue. This paper engineers that path; every result is a model projection pending measurement. First, a deconstruction-aware model: on the NVIDIA Rubin GPU the emulated rate reaches the arithmetic roof $P_{\rm FP8}/(3r+1)$ ($\approx 473$ TFLOPS at $r=12$) only within one thread-block cluster; larger outputs are re-split on the fly and held at a floor of $\approx 235$ TFLOPS (half the roof, a ratio of three design integers, not a fit), while real solvers' tall/skinny shapes stay near the crossover, $1.6$-$1.9\times$ over simple deconstruction today. Second, the method: convert-once residue workspaces, an exact two-limb constant-reduction GEMM on integer tensor pipes (or pure-SIMT dp4a), and conversion pipelined behind the MMAs, moving the crossover from $\approx 1211$ to $\approx 480$-$730$. Third, modulus co-design: all-byte and hybrid sets, two supply bounds and a carry-corrected E4M3 split of tail moduli. Fourth and central, the closed-form floor names its hardware escape, and the prize is Rubin's: a stream-side residue-conversion mode on the asynchronous copy path (Option C), a narrow fixed-function block sized as a bill of materials, takes plane formation off the arithmetic pipes and lifts the floor from 235 TFLOPS to the full 473-TFLOPS roof at unchanged cluster reach, about doubling HPL-class FP64 per Rubin GPU, and unbinds conversion-bound sparse kernels. The NVIDIA GB300 GPU, whose 135-TFLOPS roof sits at its own floor, gains little; floor and remedy are Rubin-scale. Application traces ground the analysis; constants are script-checked.

CommentsEssentially a part 3 paper of the FP8 is all you need work but also standalone work to significantly enhance the Ozaki II scheme as well as hardware assists to further accelerate FP64 emulation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑