AI 中文总结
针对大矩阵奇异值分解的难题,提出并行化快速对角化算法,可实现 m² 倍加速,在模板匹配等场景中大幅缩短运行时间,提升特征恢复效率。
AI 中文摘要
许多计算问题具有以下三个特征:(1)若能对线性算子进行对角化,该问题可高效求解;(2)因规模过大,直接对角化不可行;(3)由于问题对坐标变换具有对称性,或算子为块循环矩阵,该算子已知可与置换矩阵交换。例如,高分辨率模板匹配需要对大型高分辨率图像与详尽的相关模板列表进行重复卷积,若对相关线性算子进行对角化,可通过低秩近似对其进行压缩。所需的分解运算代价高昂,但整个问题对平面内旋转具有对称性,在此类情况下,本征函数受对称性约束,这些约束可实现高效分解。我们提出一种并行化算法,可对任意块循环矩阵进行快速对角化。当索引空间可划分为 l 类,每类包含 m 个可互换元素时,该算法可将存储成本降低 m 倍;若有 w 个工作进程可用,每个工作进程所需的浮点运算量为 O(l²m log(m)/w) + O(ml³/w),可实现 m² 倍的加速。我们使用该方法对高精度模板匹配矩阵进行分解,在简化问题上的运行时间对比显示,该快速方法每恢复一个特征的速度快 205 倍,可恢复的特征数量为原来的 22.5 倍;在典型模板匹配矩阵上,该分解所需时间比计算矩阵元素的时间少 25 倍。我们还展示了对约 30 倍大的矩阵进行分解,该矩阵涵盖了 2 Å 分辨率的冷冻电镜图像中所有可能的投影,分解仅需 14 分钟。该方法更稳定,可实现平面内旋转的任意精度恢复。
英文摘要
Many computational problems possess the following three features: (1) the problem could be solved efficiently if a linear operator could be diagonalized, (2) direct diagonalization is infeasible due to scale, but, (3) the operator is known to commute with a permutation since the problem is (a) symmetric with respect to a change of coordinates, or (b) the operator is block circulant. For example, high-resolution template matching demands repeated convolution of large, high resolution images with an exhaustive list of related templates. The associated linear operator could be compressed, via a low rank approximation, if diagonalized. The required decomposition expensive, but, the entire problem is symmetric to in-plane rotations. In these cases, the eigenfunctions are constrained by the symmetry. These constraints allow efficient decomposition. We illustrate a parallelized algorithm that allows fast diagonalization of any block-circulant matrix. When the index space can be partitioned into $l$ classes of $m$ interchangeable elements, the algorithm reduces storage costs by a factor of $m$ and, if $w$ workers are available, the algorithm demands $\mathcal{O}(l^2m\log(m)/w) + \mathcal{O}(ml^3/w)$ floating point operations per worker yielding a $m^2$ speedup. We use this procedure to decompose a high precision template matching matrix. We compare runtime on a reduced problem, where the fast procedure ran 205 times faster per recovered feature, for 22.5 times as many features. On a typical template matching matrix, the decomposition required 25 times less time than computing the matrix entries. We also demonstrate decomposition of a $\sim$30 times larger matrix that covers the complete set of possible projections represented in a cryo-EM image at 2~Å resolution in 14 minutes. This procedure is more stable and allows recovery of in-plane rotations to arbitrary precision.
Comments14 pages main text, 5 figures main text, 5 pages SI, 3 SI figures