arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用 Ozaki 方案 II 加速多精度 LU 分解

Accelerating Multiple-Precision LU Decomposition with Ozaki Scheme II

Tomonori Kouya

arXiv 2609.29554首次发表:更新:

发表机构

Otemon Gakuin University(大东文化大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于 Ozaki 方案 II 的分块部分主元 LU 分解,用整数模乘积和 CRT 重构替代多精度 GEMM,并实现直接格式转换及多种 GPU 后端,在病态系统求解中达到高达 2.84 倍加速。

AI 中文摘要

求解病态线性系统需要精度高于 binary64 的 LU 分解。标准方法——GMP/MPFR,或多分量算术如 double-double——使得 O(n^3) 更新中的每个标量乘加都处于多精度,因此无法利用当前硬件中占主导地位的低精度矩阵引擎。我们在 Ozaki 方案 II 上构建了分块、部分主元的 LU 分解,该方案用精确整数模乘积后接显式 CRT 重构替代多精度 GEMM,并在 CPU 和 GPU 上为多分量(DD/TD/QD)和任意精度实现了它。两个贡献使其实用化:非重叠扩展格式与内部定点表示之间的直接转换,省去了 MPFR 往返,带来 2.4--4.7 倍的加速;以及新的 FP16、FP8 和 binary64 GPU 后端,其中 FP8 后端使用平衡的基-17 两位编码,保持了 INT8 的完整 362.8 位 CRT 容量。在 Arm/GB10 和 x86/H100 上,最快的后端随机器而异:INT8 在 GB10 上胜出,而在 H100 上 binary64 对 QD 最快,击败了原生实现和 INT8。因此,Ozaki 方案 II 的关键点不在于使用低精度引擎,而在于选择使“每模数位数乘以引擎吞吐量”最大化的格式。对于 Lotkin 矩阵,在 p ~ 1.2 log2 cond(A) 时,我们在 n=2048 达到 10^-645 的相对误差,比完全 OpenMP 并行的多精度 LU 快达 2.84 倍。我们还通过测量表明,方案 II 相对于方案 I 的 O(p) 优势仅适用于 GEMM 项:模归约和 CRT 按 O(p^2) 增长,并在所考虑的矩阵规模下主导运行时间。该实现已开源发布。

英文摘要

Solving ill-conditioned linear systems needs LU decomposition in precisions beyond binary64. The standard approaches -- GMP/MPFR, or multi-component arithmetic such as double-double -- leave every scalar multiply-add inside the O(n^3) update in multiple precision, and so cannot exploit the low-precision matrix engines that dominate current hardware. We build a blocked, partially pivoted LU decomposition on Ozaki scheme II, which replaces the multiple-precision GEMM by exact integer modular products followed by an explicit CRT reconstruction, and implement it for multi-component (DD/TD/QD) and arbitrary precision, on CPUs and GPUs. Two contributions make this practical: a direct conversion between the non-overlapping expansion format and the internal fixed-point representation, which removes the MPFR round trip and is worth a factor of 2.4--4.7; and new FP16, FP8 and binary64 GPU back-ends, the FP8 one using a balanced base-17 two-digit encoding that keeps the full 362.8-bit CRT capacity of INT8. On an Arm/GB10 and an x86/H100 the fastest back-end changes with the machine: INT8 wins on GB10, whereas on H100 binary64 is fastest for QD, beating both the native implementation and INT8. The essential point of Ozaki scheme II is thus not to use a low-precision engine but to choose the format maximising bits-per-modulus times engine throughput. For the Lotkin matrix at p ~ 1.2 log2 cond(A) we reach relative errors of 10^-645 at n=2048, up to 2.84x faster than a fully OpenMP-parallel multiple-precision LU. We also show by measurement that the O(p) advantage of scheme II over scheme I applies only to the GEMM term: modular reduction and CRT grow as O(p^2) and dominate the runtime at the matrix sizes considered here. The implementation is released as open source.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑