AI 中文总结
该研究针对GPU上块代数多重网格的Galerkin乘积,提出共享内存分块内核及延拓算子过滤新算法,减少数据移动并缩短求解时间,用于PETSc的GPU驻留块流水线。
AI 中文摘要
Galerkin三重乘积$A_c = P^T A P$是代数多重网格(AMG)每次求解的 recurring setup 成本的主要来源。对于PDE系统的AMG,该乘积是矩形块稀疏矩阵三重乘积:三维弹性问题的精细算子具有$3\times3$块,延拓算子(prolongator)为$3\times6$,粗算子为$6\times6$,这种形状没有厂商稀疏库支持。我们在显式DRAM/L2流量模型下绘制其算法空间——经典双趟、融合重计算、重排调度、共享内存分块及 inspector-executor 变体,并使用新的PETSc块矩阵类型在可移植Kokkos(CUDA)和原生CUDA后端中实现领先变体。在NVIDIA A100上验证,该模型可预测各层级哪个变体移动的字节数最少。基于此,一种带排序无搜索调度的共享内存分块内核,在精细层级乘积中移动的字节数更少,耗时不到可移植Kokkos团队内核的一半(DRAM为10.5 vs 17.4 GB,时间为45 vs 82 ms),其完整乘积达到模型流底的$2.5$倍以内,$A\cdot P$阶段达到$1.9$倍。我们还提出延拓算子过滤(prolongator filtering),这是一种新的PETSc GAMG算法,基于Frobenius准则在保留内核的投影下从粗空间中删除小块;它减少了$P^TAP$流量、粗算子填充和内存,并在迭代次数不变的情况下将精细网格上耗时关键的$P^TAP$时间缩短了$2.9$倍。驱动应用是PETSc中完全驻留在GPU上的块流水线:有限元组装直接写入块设备矩阵,AMG setup、Galerkin乘积和求解均在主块数据上操作, recurring phases 中无标量扩展和无算子大小的设备-主机传输。
英文摘要
The Galerkin triple product $A_c = P^T A P$ dominates the recurring per-solve setup cost of algebraic multigrid (AMG). For AMG on systems of PDEs the product is a rectangular-block sparse matrix triple product: for 3D elasticity the fine operator has $3\times3$ blocks, the prolongator $3\times6$, and the coarse operator $6\times6$, a shape no vendor sparse library supports. We map its algorithm space -- classical two-pass, fused-recompute, schedule-reordered, shared-memory-tiled, and inspector-executor variants -- under an explicit DRAM/L2 traffic model, and implement the leading variants in portable Kokkos (CUDA) and native CUDA backends using new PETSc blocked matrix types. Validated on an NVIDIA A100, the model predicts per level which variant moves the fewest bytes. Guided by it, a shared-memory-tiled kernel with a sorted, search-free schedule moves fewer bytes in less than half the time of the portable Kokkos team kernels on the fine-level product (10.5 vs 17.4 GB of DRAM, 45 vs 82 ms), within $2.5\times$ of the model's streaming floor for the full product and $1.9\times$ on its $A\cdot P$ stage. We further present prolongator filtering, a new PETSc GAMG algorithm that drops small blocks from the coarse space under a Frobenius criterion with a kernel-preserving projection; it reduces $P^TAP$ traffic, coarse-operator fill, and memory, and cuts the hot $P^TAP$ time $2.9\times$ on the fine grid with iteration counts unchanged. The driving application is a fully GPU-resident blocked pipeline in PETSc: finite-element assembly writes directly into the blocked device matrix, and the AMG setup, Galerkin products, and solve all operate on primary blocked data with no scalar expansion and no operator-sized device-host transfers in the recurring phases.