arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28941math.NAcs.MScs.NA

基于隐式Householder QR分解的GPU加速分块自适应随机范围查找器

A GPU-Accelerated Blocked Adaptive Randomized Range Finder Based on an Implicit Householder QR Decomposition

  • Jülich Supercomputing Centre Forschungszentrum Jülich(于利希超级计算中心 于利希研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Carolin Penke, Andreas Herten

AI总结:

本文提出一种基于隐式Householder QR分解的GPU加速分块自适应随机范围查找器,通过混合CPU-GPU实现,可在异构系统上高效稳定地完成低秩近似,提升矩阵运算速度。

AI中文摘要:

低秩方法可降低深度神经网络训练的内存与计算需求,例如GaLore这类方法。随机范围查找器是奇异值分解的有吸引力替代方案,尤其当所需秩由规定的近似容差自适应确定时。本文提出一种基于隐式Householder QR分解的分块自适应随机范围查找器,采用优化的混合CPU-GPU实现,该方法避免了显式再正交化。数值实验表明,在无再正交化的分块Gram-Schmidt变得不稳定的场景中,此方法能保持正交性与近似精度。分块形式可暴露矩阵-矩阵操作,实现CPU面板分解与GPU更新的重叠。在NVIDIA GH200上,重叠实现将测试的最大矩阵运行时间从CPU的9.91秒降至0.407秒,为异构系统上的低秩近似提供了稳定高效的构建块,可应用于计算科学与工程领域。

英文摘要:

Low-rank methods can reduce the memory and computational requirements of deep neural network training in approaches such as GaLore. Randomized range finders offer an attractive alternative to singular value decompositions, particularly when the required rank is determined adaptively from a prescribed approximation tolerance. We introduce a blocked adaptive randomized range finder based on an implicit Householder QR decomposition and an optimized hybrid CPU--GPU implementation. The proposed method avoids explicit reorthogonalization. Numerical experiments show that it preserves orthogonality and approximation accuracy in regimes where block Gram--Schmidt without reorthogonalization becomes unstable. The blocked formulation exposes matrix--matrix operations and enables overlap of CPU panel factorization with GPU updates. On an NVIDIA GH200, the overlapped implementation reduces the runtime for the largest tested matrix from 9.91 seconds on the CPU to 0.407 seconds. The method provides a stable and efficient building block for low-rank approximation on heterogeneous systems with applications in computational science and engineering.

↑