发表机构
NVIDIA Corporation(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对TME模型补充了fp64仿真的输入解构代价项,推导了运算强度阈值,修正了原论文对部分内核性能的错误估计,并给出了预计算残差的相关结论。
AI 中文摘要
“FP8 is All You Need(第一部分)”中的张量内存平衡(TME)模型,将Ozaki方案II的fp64仿真执行时间计算为张量核心项、高带宽内存(HBM)流量项与每个输出重构项的最大值。但该模型遗漏了每个输入的解构代价:在任何矩阵乘法执行前,每个流式fp64操作数必须在SIMT管道上针对r个模中的每个进行缩放、舍入和模约简。本文补充了这第四项,从cuBLAS仿真路径校准其常数,并推导了闭式运算强度阈值OI* = c_q r P_fp64/(8P_int),低于该阈值时,无论张量核心吞吐量如何,仿真都无法匹配原生fp64性能。在NVIDIA B300 GPU上,该阈值约为0.56 FLOP/B。结果显示,原论文声称可加速的内存受限内核(GEMV、SpMV及低批量GEMV)性能被限制为原生性能的0.3-0.9倍,7点模板的性能被限制为1.8倍,而非原论文声称的3.1倍;稠密GEMM则如预期般不受影响。本文还表明,预计算并存储残差会将相同代价转移至带宽项,并给出了实现需达到的指令数以推翻该界限。
英文摘要
The Tensor-Memory Equilibrium (TME) model of "FP8 is All You Need (Part 1)" calculates the execution time of Ozaki Scheme II emulation of fp64 as the maximum of a tensor-core term and a High-Bandwidth Memory (HBM) traffic term, plus a per-output reconstruction term. However, it omits the per-input deconstruction cost: every streamed fp64 operand must be scaled, rounded, and reduced modulo each of the $r$ moduli on SIMT pipes before any matrix multiply can issue. In this note we add this fourth term, calibrate its constant from the cuBLAS emulation path, and derive a closed-form operational-intensity threshold $\mathrm{OI}^{*} = c_q r P_{\mathrm{fp64}}/(8P_{\mathrm{int}})$ below which emulation cannot match native fp64 regardless of tensor-core throughput. On the NVIDIA B300 GPU the threshold is $\mathrm{OI}^{*}\approx 0.56$ FLOP/B. As a result, GEMV, SpMV, and low-batch GEMV, which are the memory-bound kernels the original paper claims to accelerate, are limited to 0.3-0.9x of native performance, and the 7-point stencil to 1.8x rather than the claimed 3.1x. Dense GEMM is unaffected as expected. We also show that precomputing and storing the residues moves the same cost into the bandwidth term, and we state the instruction count that an implementation would have to achieve to invalidate the bound.