发表机构
CEA; Aix-Marseille University(法国原子能与替代能源委员会; 艾克斯-马赛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对 GPU 上大量微型线性系统求解难题,本文比较多种 LUpp 求解器并设计专用 CUDA 内核,在本构定律评估中实现高达 6.5 倍(对 cuSolverDx)和 17.7 倍(对 MAGMA)的加速。
AI 中文摘要
许多应用需要在 GPU 上求解大量独立的线性系统。虽然这一需求对于小型到大型系统已得到很好的解决,但微型系统(此处指维度低于 32 的系统)仍然具有挑战性。这在本构定律评估中尤其重要,其中数百万个积分点被独立处理,且每次本构更新通常依赖于牛顿迭代法。每次迭代都需要使用带部分主元的 LU 分解(LUpp)来求解一个双精度微型一般稠密线性系统。本研究在一个闭源原型中进行,该原型作为将 TFEL/MFront(一个用于材料知识的开源代码生成工具)目前在 CPU 上提供的本构定律评估移植到 NVIDIA GPU 的演示器。我们比较了多种双精度 LUpp 求解器,包括来自 GPU 线性代数库的实现以及自定义设计的 CUDA 内核。比较首先在大型独立线性系统批次上进行,然后在完整的本构定律评估工作流中进行,其中每个积分点需要一系列不同的线性系统,每个系统对应其自身牛顿循环的一次迭代。研究表明,最佳的 LUpp 求解策略强烈依赖于多个因素,包括系统大小和应用上下文。我们讨论了几个关键方面,包括寄存器压力、占用率、直接从设备代码调用 LUpp 求解器的能力,以及为每个系统分配多个线程是否是最有效的策略。在 NVIDIA H100 GPU 上的实验表明,本文提出的专用 LUpp 求解器在这类工作负载上可以优于现有的最先进方法,相对于 cuSolverDx 加速高达 6.5 倍,相对于 MAGMA 加速高达 17.7 倍。
英文摘要
Many applications require solving large numbers of independent linear systems on GPUs. While this need is well addressed for small to large systems, tiny ones, understood here as systems of dimension below 32, remain challenging. This is especially relevant in constitutive law evaluation, where millions of integration points are handled independently, and where each constitutive update generally relies on a Newton iterative method. Each iteration then requires the double-precision solution of a tiny general square linear system using LU factorization with partial pivoting (LUpp). The present study was conducted within a closed-source prototype, which serves as a demonstrator for porting to NVIDIA GPUs constitutive law evaluations currently provided on CPUs by TFEL/MFront, an open-source code generation tool for material knowledge. We compare several double-precision LUpp solvers, including implementations from GPU linear algebra libraries as well as custom-designed CUDA kernels. The comparison is performed first on large batches of standalone linear systems, and then within the full constitutive-law evaluation workflow, where each integration point requires a sequence of distinct linear systems, one per iteration of its own Newton loop. The study shows that the best LUpp solving strategy strongly depends on several factors including system size and application context. We discuss several key aspects, including register pressure, occupancy, the ability to invoke the LUpp solver directly from device code, and whether assigning several threads to each system is the most efficient strategy. Experiments on an NVIDIA H100 GPU show that the specialized LUpp solvers proposed in this work can outperform existing state-of-the-art approaches for this class of workloads, with speedups of up to 6.5x over cuSolverDx, and up to 17.7x over MAGMA.