发表机构
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究为经典Kaczmarz方法建立了紧的线性和次线性收敛界,以矩阵行范数、相关性等参数解释其高效收敛,并证明循环更新在弱相关行时优于随机更新。
AI 中文摘要
Kaczmarz经典方法于1937年提出,是求解线性方程组$A x = b$的教科书式迭代方法。尽管该方法被广泛使用,特别是在求解反问题的背景下,通常包含在Matlab、Python和Julia的软件包中,但其收敛性的精确刻画长期以来被认为难以获得。虽然已经建立了不同的收敛速率界,但这些界在很大程度上不能令人满意,因为它们无法解释经典循环更新方法在实践中高效收敛的现象。在这项工作中,我们获得了经典Kaczmarz方法收敛性的紧刻画,建立了线性和次线性收敛界。最坏情况下的紧界(即,在所考虑族中的某些实例上精确达到)以仅依赖于$A$的固定矩阵表示,但无法完全用矩阵谱和行相关性来解释,而这些已被观察到对收敛有影响。因此,我们提供了这些界的松弛形式,这些松弛界完全可以用矩阵行范数、行相关性、秩和极值正奇异值来表示。所提供的松弛界解释了在矩阵行全部相互平行或正交的特殊情况下的单周期收敛。我们进一步论证,界中出现的不同参数的依赖性是必要的,并且在最坏情况下与可达到的最佳值相差一个小的常数因子。最后,我们的界解释了为什么当矩阵行弱相关时,经典循环更新比随机更新更快,这在通常使用循环方法的反问题中经常观察到。
英文摘要
The classical method of Kaczmarz, introduced in 1937, is a textbook iterative method for solving linear systems $A x = b.$ Despite its widespread use, particularly in the context of solving inverse problems where it is commonly included in software packages within Matlab, Python, and Julia, the precise characterization of convergence has long been deemed difficult to obtain. While different bounds on convergence rates have been established, they are largely unsatisfying as they cannot explain the classical, cyclic-update method's efficient convergence in practice. In this work, we obtain a tight characterization of convergence of the classical Kaczmarz method, establishing both linear and sublinear convergence bounds. The worst-case tight (i.e., exactly attained by some instances in the considered family) bounds are expressed in terms of a fixed matrix that depends only on $A,$ but are not fully interpretable in terms of the matrix spectrum and row correlations, which had been observed to have an impact on convergence. We thus provide relaxations of these bounds that are fully expressible in terms of matrix row norms, row correlations, rank, and extremal positive singular values. The provided relaxed bounds explain one-cycle convergence in special cases where the matrix rows are all either parallel or orthogonal to each other. We further argue that the dependence on different parameters appearing in the bounds is necessary and within a small constant factor of the best attainable in the worst case. Finally, our bounds explain why the classical cyclic update is faster than the randomized one when matrix rows are weakly correlated, which is often observed in inverse problems where the cyclic method is used.