AI 中文总结
本文提出HSMA-TRSM优化框架,针对GPU平台的TRSM,通过分层共享内存感知设计,在三类加速器上相较cuBLAS、rocBLAS获最高2.05倍、2.06倍加速比。
AI 中文摘要
多右端项三角求解(TRSM)是支撑LU分解、Cholesky分解、稀疏直接求解器及矩阵求逆的基础BLAS Level-3操作。本文研究左侧下三角情形,其高效GPU实现仍具挑战性,因前向替换引入严格行依赖,且共享内存过稀缺,无法容纳双复数等宽数据类型的两个操作数矩阵。本文提出HSMA-TRSM,即面向NVIDIA A100、NVIDIA H800及海光DCU Z100加速器的左侧下三角TRSM的分层共享内存感知优化框架。针对小规模场景(m,n≤64),通过循环展开与指令重排设计流水线式计算-内存重叠机制,并提出双线程组七阶段流水线策略以解决双复数类型的共享内存约束;针对大规模问题,引入对角块解耦优化,其对角块求逆的共享内存占用为O(IB),可基于矩阵规模与硬件特性实现自适应块大小选择。基于离线分析与在线查找的编译时配置选择框架为各平台选择最优块大小,零运行时开销。在NVIDIA A100、H800及海光DCU Z100上评估显示,HSMA-TRSM相较cuBLAS实现峰值加速比2.05倍,相较rocBLAS实现峰值加速比2.06倍。该增益在受共享内存约束的双复数小规模案例、自适应块优化改善GEMM主导更新的大规模实数案例中最为显著,而成熟的厂商内核在部分场景中预留的优化空间较小。
英文摘要
Triangular Solve with Multiple Right-hand Sides (TRSM) is a fundamental BLAS Level-3 operation that underpins LU/Cholesky decomposition, sparse direct solvers, and matrix inversion. In the left-side lower-triangular case studied in this paper, efficient GPU implementation remains challenging because forward substitution introduces strict row-wise dependencies, and shared memory is too scarce to hold both operand matrices for wide data types such as double complex. This paper presents HSMA-TRSM, a hierarchical shared memory-aware optimization framework for left-side lower-triangular TRSM on NVIDIA A100, NVIDIA H800, and Hygon DCU Z100 accelerators. For the small-scale regime (m,n<=64), we design a pipelined compute-memory overlap mechanism through loop unrolling and instruction reordering, and propose a dual thread-group seven-stage pipeline strategy to address shared memory constraints for double complex types. For large-scale problems, we introduce a diagonal block decoupling optimization with an O(IB)shared-memory footprint for diagonal block inversion, enabling adaptive block size selection based on matrix scale and hardware characteristics. A compile-time configuration selection framework based on offline profiling and online lookup selects the optimal block size per platform with zero runtime overhead. Evaluated on NVIDIA A100, H800, and Hygon DCU Z100, HSMA-TRSM achieves peak speedups of 2.05xover cuBLAS and 2.06xover rocBLAS. The gains are strongest in shared-memory-constrained double-complex small cases and in large real-type cases where adaptive blocking improves GEMM-dominated updates, while mature vendor kernels leave less optimization headroom in some regimes.
CommentsAccepted for publication in the Proceedings of the 55th International Conference on Parallel Processing (ICPP '26), 2026. To appear in the ACM Digital Library. 11 pages