AI 中文总结
该研究提出通过求解多个并行化的线性系统副本,结合块共轭梯度法,在不额外增加硬件资源的情况下,将稠密系统求解时间最多缩短6倍、稀疏系统最多缩短5倍,且可兼容各类预条件子。
AI 中文摘要
我们提出一种在并行计算硬件上加速共轭梯度法(CG)的方法,通过对一个线性系统执行块共轭梯度来实现。我们的目标是通过求解该系统的多个副本,同时探索解空间的多个区域,以减少收敛所需的迭代次数。尽管求解多个系统比求解单个系统需要更多的浮点运算,但这些计算具有高度并行性,即使无需额外硬件资源也可实现。我们的方法不取代预处理,可与任何预条件子配合使用。我们开发了线性回归模型,以系数矩阵的规模、非零元素数量以及CG求解时间为函数,预测迭代减少量和加速比。实验表明,对于规模较大且CG求解时间较长的稠密系统,我们的方法可将求解时间最多减少6倍;对于稀疏系统,最多可减少5倍。
英文摘要
We propose a method to accelerate the conjugate gradient method (CG) on parallel computing hardware by performing block conjugate gradient on one linear system. We aim to reduce the number of iterations until convergence by solving several copies of the system to explore multiple regions of the solution space at once. Although solving multiple systems requires more floating point operations than solving one, these computations are highly parallelizable, even without additional hardware resources. Our method does not replace preconditioning and can be used with any preconditioner. We developed linear regression models to predict iteration reduction and speedup as functions of the coefficient matrix's size, its number of nonzeros, and the solve time with CG. Our experiments show that our method can reduce solve time by up to 6 times in dense systems and up to 5 times in sparse systems that require long solve times with CG relative to their size.
Comments17 pages, 3 figures