发表机构
University of California Berkeley; Rice University; University of Chicago(加州大学伯克利分校; 莱斯大学; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出算术代价最优的并行雅可比算法,针对对称特征值问题和SVD,在二维、2.5维处理器网格下分析其算术与通信复杂度,扩展结果至单边雅可比SVD。
AI 中文摘要
本文针对对称特征值问题和奇异值分解(SVD),提出了雅可比方法的若干并行版本。作为[Demmel、Luo、Schneider和Wang 2025]工作的延续,我们开发了算术代价最优的并行雅可比算法,其带宽或延迟可匹配并行矩阵乘法的相应下界。我们的研究聚焦于具有可变处理器布局的标准分布式内存环境,包括二维和2.5维处理器网格。在二维情形下,我们证明并行雅可比的标准实现可实现算术代价的完美加速——即使用P个处理器时,复杂度为O(n³/P),同时达到二维矩阵乘法的带宽下界,且(几乎)达到延迟下界。通过采用2.5维处理器网格并利用2.5维矩阵乘法,等效于增加每个处理器的内存,我们证明并行雅可比可实现更低的带宽/延迟,但我们也证明,在任何雅可比算法中,这些代价无法同时匹配并行矩阵乘法的已知最优下界。最后,我们将结果扩展到单边雅可比SVD。
英文摘要
This paper presents several parallel versions of Jacobi's method for the symmetric eigenvalue problem and the SVD. A continuation of [Demmel, Luo, Schneider, & Wang 2025], we develop parallel Jacobi algorithms whose arithmetic cost is optimal and whose bandwidth or latency can match the corresponding lower bounds of parallel matrix multiplication. Our focus is a standard distributed-memory setting with variable processor layouts, including both 2D and 2.5D processor grids. In the 2D case, we demonstrate that a standard implementation of parallel Jacobi achieves a perfect speedup in arithmetic cost -- i.e., complexity $O(n^3/P)$ when done with $P$ processors -- while hitting the 2D matrix-multiplication lower bound for bandwidth and (nearly) the lower bound for latency. By employing a 2.5D processor grid and leveraging 2.5D matrix multiplication, equivalently by increasing the memory per processor, we demonstrate that parallel Jacobi can achieve even lower bandwidth/latency, though we also prove that these costs cannot simultaneously match the best-known bounds for parallel matrix multiplication in any Jacobi algorithm. Finally, we extend our results to one-sided Jacobi SVD.
Comments42 pages, 6 figues, 2 tables