用于平面波密度泛函理论的分块随机Davidson本征求解器
Randomized Block Davidson Eigensolvers for Plane-Wave Density-Functional Theory
浏览论文内容
中文总结 AI 辅助
本文提出分块随机Davidson本征求解器,用随机Gram-Schmidt替代正交化,在特定矩阵规模下比确定性求解器更快,其收益受所需态与问题维度缩放关系影响。
中文摘要 AI 辅助
迭代对角化是平面波密度泛函理论(DFT)的主要开销,其中搜索空间正交化的计算成本随问题规模和目标态数量的增长速度极快。本文提出一种分块随机Davidson型本征求解器,该求解器用带草图内积的随机Gram-Schmidt方法替代欧几里得正交化,仅需对基进行一次遍历,同时保持其条件数与输入向量无关。该修改仅改变了Rayleigh-Ritz步骤,使其成为一个确定的广义Hermitian本征问题。Ritz提取保持精确,保留了真实Ritz对和交错性质,该性质使得每个能带能量都是真实能量的上界。该方法通过单一Julia代码实现了CPU和GPU的混合精度,与DFTK进行无矩阵接口,并在开源RandESC库中发布。在固定本征对数量的稀疏测试问题中,当矩阵尺寸超过约2×10^4时,带草图的求解器超过其确定性对应求解器,在5×10^5的矩阵尺寸下快25%。然而,在完全自洽场DFT计算中,两种Davidson变体仅比局部最优分块预条件共轭梯度(LOBPCG)参考方法的总时间快5%至11%,而草图的额外收益有限。随着所需态数量随系统规模增长,正交化的节省被广义本征问题抵消。因此,草图发挥作用的范围由所需态数量与问题维度的缩放关系决定,而非本征求解器本身。
英文摘要
Iterative diagonalization is the dominant cost of plane-wave density-functional theory (DFT), with search-space orthogonalization scaling particularly quickly with problem size and the number of target states. We present a randomized block Davidson-type eigensolver that replaces Euclidean orthogonalization with randomized Gram-Schmidt in a sketched inner product, requiring only a single pass over the basis while keeping its conditioning bounded independently of the input vectors. This modification changes only the Rayleigh-Ritz step, which becomes a definite generalized Hermitian eigenproblem. Ritz extraction remains exact, preserving true Ritz pairs and the interlacing property that makes each band energy an upper bound on the true one. The method is implemented in mixed precision for CPUs and GPUs from a single Julia code, interfaces matrix-free with DFTK, and is released in the open-source RandESC library. On sparse test problems with a fixed number of eigenpairs, the sketched solver overtakes its deterministic counterpart beyond matrix dimensions of about $2\times 10^4$ and is $25\%$ faster at $5\times 10^5$. In full self-consistent field DFT calculations, however, both Davidson variants outperform the locally optimal block preconditioned conjugate gradient (LOBPCG) reference only by $5$ to $11\%$ in total time, while the additional benefit of sketching is limited. As the number of requested states grows with system size, orthogonalization savings are offset by the generalized eigenproblem. Therefore, the regime in which sketching pays off is set by how the number of wanted states scales with the problem dimension, not by the eigensolver as such.