fp8 中的快速矩阵乘法:认证系数优化与实测误差
Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error
浏览论文内容
中文总结 AI 辅助
本文提出系数泛函 Phi 最小化方法,在 fp8 下认证 Strassen 型矩阵乘法实现的全局最优性,并实测其误差优势,将算法实现转化为数学认证的设计问题。
中文摘要 AI 辅助
Strassen 型算法有许多实现,它们具有相同的精确乘积和乘法次数,但由于基变换会重塑系数几何结构,因此 fp8 误差不同,这就提出了运行哪一种实现的问题。目前没有任何理论能解决这个问题:经典稳定性控制的是最坏情况下的 $\ell_1$ 增长,而非期望误差幅度,而且 Dumas--Pernet--Sedoglavic 优化器只能被称为可能最优,其全局最优性尚未得到证明。为了解决这个问题,我们为每个实现附加一个系数泛函 $\Phi$,它是其系数几何结构的标量摘要,我们在基变换轨道上将其最小化。这个 Hadamard 流形上的 Kempf--Ness 问题使我们能够证明全局 $\Phi$ 最优性,而不仅仅是搜索它:精确的矩映射零点固定了 $\Phi_{\min} = 200/9$,并且 de Groote 的分类将该最优性扩展到每一个精确实秩 7 的 $2\times2$ 分解。因此,每一个精确实秩 7 的实现,在固定噪声系数下,其 $\Phi$ 预测的 RMS 常数至少是三次算法(cubic algorithm)的 $5/3$ 倍。然后,我们引入一个显式的块缩放 e4m3 模型,其中 $\Phi$ 是相对期望均方误差的首阶系数,并针对实际 fp8 误差测试了由此产生的 $\Phi$ 预测排序。排序和重新基化实验在测试的融合块缩放范围内支持该预测,并且在来自四个架构族的真实 matmul 块上,$\Phi$ 最优实现落在 fp8 低误差区域内。在真实 deep_gemm 内核上的两个约 70B 模型中,相同的实现消除了经典 Strassen 相对于干净模型的多余 NLL 的 10% 到 55%。因此,算法实现成为一个数学上认证的设计问题,而非调优选择:一个独立的低精度轴,具有全局 $\Phi$ 最优性和实测的 fp8 相关性。
英文摘要
A Strassen-type algorithm has many realizations with the same exact product and multiplication count yet different fp8 error because basis changes reshape coefficient geometry, posing the question of which to run. No current account settles this: classical stability controls worst-case $\ell_1$ growth, not the expected-error magnitude, and the Dumas--Pernet--Sedoglavic optimizer could only be called probably optimal, its global optimality unproved. To settle this, we attach to each realization a coefficient functional $Φ$, a scalar summary of its coefficient geometry, which we minimize over the change-of-basis orbit. This Kempf--Ness problem on a Hadamard manifold lets us certify the global $Φ$ optimum rather than merely search for it: an exact moment-map zero fixes $Φ_{\min} = 200/9$, and de Groote's classification extends that optimality to every exact real rank-7 $2\times2$ decomposition. Every exact real rank-7 realization therefore has a $Φ$-predicted RMS constant at least $5/3$ times that of the cubic algorithm, at fixed noise coefficient. We then introduce an explicit block-scaled e4m3 model in which $Φ$ is the leading-order coefficient of relative expected mean-squared error, and we test the resulting $Φ$-predicted ordering against realized fp8 error. Ordering and re-basing experiments support that prediction within tested fused block-scaled regimes, and on real matmul tiles from four architecture families the $Φ$-optimal realization falls in the fp8 low-error region. Across two $\sim$70B models on real deep_gemm kernels, the same realization removes 10 to 55% of classic Strassen's excess NLL over the clean model. Algorithm realization thus becomes a mathematically certified design problem rather than a tuning choice: an independent low-precision axis with a global $Φ$ optimum and measured fp8 relevance.
发表机构
- Beijing Tongming Lake Information Technology Application Innovation Center (TLAIC)(北京通明湖信息技术应用创新中心)
- Fudan University Institute of Systems for Advanced Computing(复旦大学先进计算系统研究院)
- Harbin Institute of Technology(哈尔滨工业大学)
- Key Lab of HCST (PKU), MOE(北京大学高可信软件技术教育部重点实验室;北京大学计算机学院)
- SCS, Peking University(上海开放计算系统研究院)
- Shanghai Institute of Systems for Open Computing
机构由 AI 辅助整理,请以论文原文为准。