你只需转换两次:Ozaki 方案 II 用于链式张量模乘积
You Only Convert Twice: Ozaki Scheme II for Chained Tensor Mode Products
浏览论文内容
中文总结 AI 辅助
本工作提出 OzII-RescaleBE 方法,通过直接在余数表示上进行缩放和基扩展,实现 Ozaki 方案 II 下链式张量模乘积的持续执行,避免了中间浮点转换,在保持高精度的同时提升了吞吐量。
中文摘要 AI 辅助
本工作提出了 OzII-RescaleBE,一种在 Ozaki 方案 II 中实现链式张量模乘积持续执行的方法,将中间张量保持为余数形式。通过直接在余数表示上进行缩放和基扩展,该方法仅在链的起始处需要转换为余数,在链的末端才进行重构。提供了使用 CuTe DSL 和 INT8 张量核心的实现用于验证和评估。该方法与 Ozaki 方案 II 的逐模组合(GEMMul8-Composed)以及稠密 Kronecker 乘积公式(GEMMul8-KRON)进行了比较,两者均使用 GEMMul8 库实现。在三个具有不同输入分布的数据集上,对于链深度 d=3 至 7,OzII-RescaleBE 实现了范数相对误差范围从 2.5×10^-16 到 2.6×10^-15。相应的误差范围对于 GEMMul8-KRON 为 10^-16 到 10^-15,对于 GEMMul8-Composed 和 FP64 链为 1×10^-16 到 3×10^-16。在 d=8 时,OzII-RescaleBE 的误差增加到 7×10^-14 到 6×10^-13 之间。对于模式大小 n=8 到 256,OzII-RescaleBE 实现了吞吐量范围从 1.47 到 4.42 GDoF/s。在 n=8 和 12 时,其吞吐量在 GEMMul8-KRON 的 12% 以内,在更大的测试尺寸下超过它。与 GEMMul8-Composed 相比,OzII-RescaleBE 在 n≤128 时实现了 2.1 到 60 倍的吞吐量,但在 n=256 时吞吐量低 6 到 12%。在测量的内存比较中,它使用的设备内存也少于 GEMMul8 基线。这些结果表明,缩放和基扩展能够在避免重复中间浮点转换的同时,实现准确的余数域张量链。
英文摘要
This work proposes OzII-RescaleBE, a method that enables persistent execution of chained tensor mode products in Ozaki scheme II, keeping the intermediate tensors in residue form. By performing rescaling and base extension directly on the residue representation, it requires conversion to residues only at the beginning of the chain and reconstruction only at the end. An implementation using CuTe DSL and INT8 Tensor Cores is provided for verification and evaluation. The method is compared with per-mode composition of Ozaki scheme~II (GEMMul8-Composed) and a dense Kronecker-product formulation (GEMMul8-KRON), both implemented using the GEMMul8 library. Across three datasets with different input distributions, OzII-RescaleBE achieves normwise relative errors ranging from $2.5\times10^{-16}$ to $2.6\times10^{-15}$ for chain depths $d=3$ through $7$. The corresponding errors range from $10^{-16}$ to $10^{-15}$ for GEMMul8-KRON and from $1\times10^{-16}$ to $3\times10^{-16}$ for GEMMul8-Composed and the FP64 chain. At $d=8$, the errors of OzII-RescaleBE increase to between $7\times10^{-14}$ and $6\times10^{-13}$. For mode sizes $n=8$ to $256$, OzII-RescaleBE achieves throughputs ranging from 1.47 to 4.42\,GDoF/s. Its throughput is within 12\,\% of GEMMul8-KRON at $n=8$ and $12$ and exceeds it at larger tested sizes. Compared with GEMMul8-Composed, OzII-RescaleBE achieves $2.1$--$60\times$ the throughput for $n\leq128$, but has 6--12\,\% lower throughput at $n=256$. It also uses less device memory than the GEMMul8 baselines in the measured memory comparisons. These results demonstrate that rescaling and base extension enable accurate residue-domain tensor chains while avoiding repeated intermediate conversions to floating point.
发表机构
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。