arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稠密矩阵皆相似;稀疏矩阵各有各的稀疏:一种结构自适应的分块Cholesky分解

Dense Matrices Are Alike; Sparse Matrices Are Sparse in Their Own Way: A Structure-Adaptive Tile Cholesky Factorization

Esmail Abdul Fattah, Hatem Ltaief, Håvard Rue, David E. Keyes

arXiv 2609.29765首次发表:更新:

发表机构

King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种结构自适应的分块Cholesky分解,根据稀疏模式动态选择稠密、稀疏或半稀疏存储,在60个SPD矩阵上比固定结构快1.6-12.5倍,并首次扩展到GPU。

AI 中文摘要

稀疏直接Cholesky求解器为整个矩阵固定一种数据结构,但对称正定系统范围从近乎稠密到不规则,有时在一个矩阵内混合两者。我们让数据结构跟随稀疏结构,跨矩阵以及矩阵内的分块。在数值工作开始之前,一个轻量级选择器捕获Cholesky因子的稀疏模式,并在一个静态共享内存调度上将矩阵路由到三种模式之一:稠密、稀疏或中间半稀疏模式。稠密分块以完整形式存储;半稀疏模式使用一种新的活动列分块,仅保留分解将触及的列,因此空间模型中常见的带状或箭头状结构仍能达到BLAS-3效率。该方法适用于集成嵌套拉普拉斯近似(INLA):一种稀疏模式被分解数千次,每次数值不同,因此一次性分析成本被摊销,每次分解节省累积。我们在60个SPD矩阵上评估该技术,在Intel Xeon和AMD EPYC节点上与MUMPS、PaStiX、CHOLMOD、symPACK和Intel oneMKL PARDISO进行比较;仅选择器就在每种模式下实现了最低的总分解时间。在整个测试集上,它比最佳固定单一结构模式快1.6到2.6倍,在Intel上比所有替代方案快1.8到12.5倍,在AMD上快2.6到10.1倍,最大增益出现在最昂贵的分解上。它用更多的一次性分析换取每次分解更少的时间,在给定模式的第三次分解时拉开差距。作为首个GPU扩展,在一个NVIDIA A100上,因子驻留在设备上,稠密路径比更快的CPU节点快1.2到6.3倍,差距随因子大小增大而扩大。求解器、Python/R/Julia接口、基准测试套件和结果在此https URL开放。

英文摘要

Sparse direct Cholesky solvers fix one data structure for an entire matrix, but symmetric positive definite systems range from nearly dense to irregular, sometimes mixing both within one matrix. We let the data structure follow the sparsity structure, across matrices and across tiles within a matrix. Before numerical work starts, a lightweight selector captures the sparsity pattern of the Cholesky factor and routes the matrix, on one static shared-memory schedule, to one of three regimes: dense, sparse, or an intermediate semisparse regime. Dense tiles are stored in full; the semisparse regime uses a new active-column tile that keeps only the columns the factorization will touch, so banded or arrowhead-shaped structures common in spatial models still reach BLAS-3 efficiency. The approach suits the integrated nested Laplace approximation (INLA): one sparsity pattern factorized thousands of times with different values, so the one-time analysis cost is amortized and every factorization saving compounds. We evaluate the technique on 60 SPD matrices, comparing against MUMPS, PaStiX, CHOLMOD, symPACK, and Intel oneMKL PARDISO on Intel Xeon and AMD EPYC nodes; the selector alone achieves the lowest total factorization time in every regime. Summed over the suite, it beats the best fixed single-structure mode by 1.6 to 2.6x, and every alternative by 1.8 to 12.5x on Intel and 2.6 to 10.1x on AMD, with the largest gains on the most expensive factorizations. It trades more one-time analysis for less time per factorization, pulling ahead by the third factorization of a given pattern. As a first GPU extension, the dense route on one NVIDIA A100, with the factor resident on the device, runs 1.2 to 6.3x faster than on the faster CPU node, the margin widening with factor size. Solver, Python/R/Julia interfaces, benchmark suite, and results are open at https://github.com/esmail-abdulfattah/sTiles.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑