arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06414math.NAcs.CEcs.NA

矩阵分裂方法的混合精度模型

A Mixed-Precision Model for Matrix-Splitting Methods

Neil Lindquist, David Appelhans, Chales W. Jackson, Joseph M. Derlaga

首次发表
浏览论文内容

中文总结 AI 辅助

针对内存受限的平稳迭代方法,提出混合精度技术以减少数据移动,理论及实验证明收敛速率与精确精度相当,结合迭代细化可达双精度,混合半/双精度实现雅可比、高斯-赛德尔及SSOR迭代的显著加速。

中文摘要 AI 辅助

由于算术强度低,平稳迭代方法的性能受内存限制。因此,我们提出一种混合精度技术,以减少矩阵分裂迭代方法中的数据移动量。我们从理论和实验两方面证明,该技术通常能以与精确精度迭代相似的速率收敛,具体取决于所用精度及$M$矩阵的条件数。此外,我们证明可以使用迭代细化将解细化至完全双精度精度。采用这种方法,我们将半精度和双精度混合使用,与均匀双精度相比,标量和块状雅可比迭代平均加速$1.60\ imes$,标量和块状高斯-赛德尔迭代平均加速$1.25\ imes$。此外,我们降低了OVERFLOW(NASA用于可压缩流体动力学模拟的代码)中使用的SSOR迭代的精度,并为整个应用实现了$1.14\ imes$的加速。

英文摘要

The performance of stationary iterative methods is memory-bound, due to the low arithmetic intensity. So, we propose a mixed-precision technique to reduce the amount of data movement in matrix-splitting iterative methods. We demonstrate both theoretically and experimentally that this technique will generally converge at a similar rate as an exact-precision iteration, depending on the precision used and the conditioning of the $M$ matrix. Furthermore, we demonstrate that iterative refinement can be used to refine the solution to full double-precision accuracy. Using this approach, we mixed half and double precision to achieve an average speedup over uniform double precision of $1.60\times$ for scalar and block Jacobi and $1.25\times$ for scalar and block Gauss-Seidel. Furthermore, we reduced the precision in the SSOR iteration used by OVERFLOW, a NASA code for compressible fluid dynamics simulations, and achieved a $1.14\times$ speedup to the overall application.

发表机构

  • NVIDIA(英伟达)
  • NASA Langley Research Center(美国国家航空航天局兰利研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑