arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07354physics.chem-ph

超越沿用四十年的DIIS默认方案:自洽场计算的辅助曲率加速方法

Beyond the Four-Decade DIIS Default:Auxiliary-Curvature Acceleration of Self-Consistent-Field Calculations

Peng Bao, Qiang Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AURORA-SCF方法,采用辅助曲率技术,在不改变目标定态方程的前提下,通过16个CPU对和63个GPU对的实验,实现SCF计算墙钟时间26.2%-34%的缩短,超越沿用四十年的DIIS默认方案。

中文摘要 AI 辅助

直接迭代子空间反演(DIIS)自1980年提出以来,仍是加速Hartree-Fock(HF)和Kohn-Sham自洽场(SCF)计算的实用默认方法。后续诸多方法虽在鲁棒性或迭代次数上有所改进,但DIIS近乎为零的开销使其难以大幅缩短总 wall time(墙钟时间)。本文提出AURORA-SCF(辅助曲率统一黎曼轨道响应加速),该方法仅用所需目标哈密顿量计算能量和梯度,同时从独立的价层STO-3G模型获取大部分轨道曲率;目标级传输的L-BFGS历史可修正模型不匹配,且通过移位无矩阵求解、信赖域及测地线轨道更新控制步长。在16个直接CPU RHF对(覆盖137-1484个原子轨道)中,AURORA-SCF在所有案例中均更快,平均墙钟时间缩短26.2%,目标J/K构建缩短33.3%;在另一组63个收敛、能量匹配的密度拟合GPU对(覆盖7种分子、闭壳层与开壳层形式,以及HF、PBE、B3LYP和M06-2X方法)中,其在所有纳入对中仍更快,平均墙钟时间缩短26.5%,目标J/K构建缩短31.9%;聚焦的直接CPU和GPU扫描显示平均墙钟时间缩短30%-34%。这些结果证明了超越沿用四十年的DIIS默认方案的具体进展:传输的、正割修正的曲率可在不改变目标定态方程的前提下,减少总SCF墙钟时间,而非仅减少迭代次数;该优化模式还为量子化学其他领域的更快轨道优化及科学计算中的多保真度优化提供了途径。

英文摘要

Direct inversion in the iterative subspace (DIIS), introduced in 1980, remains the practical default for accelerating Hartree-Fock and Kohn-Sham self-consistent-field (SCF) calculations. Many later methods improve robustness or iteration count, but the near-zero overhead of DIIS makes a broad reduction in total wall time difficult. We introduce AURORA-SCF (auxiliary-curvature unified Riemannian orbital-response acceleration), which evaluates energy and gradient only with the requested target Hamiltonian while obtaining most orbital curvature from an independent valence-STO-3G model. A target-level, transported L-BFGS history corrects the model mismatch; a shifted matrix-free solve, trust region, and geodesic orbital update control the step. In 16 direct CPU RHF pairs spanning 137-1484 atomic orbitals, AURORA-SCF was faster in every case, reducing mean wall time by 26.2% and target J/K builds by 33.3%. In a separate set of 63 converged, energy-matched density-fitted GPU pairs spanning seven molecules, closed- and open-shell formalisms, and HF, PBE, B3LYP, and M06-2X, it was again faster in every included pair, with mean reductions of 26.5% in wall time and 31.9% in target J/K builds. Focused direct CPU and GPU sweeps show mean wall-time reductions of 30-34%. These results establish a specific advance beyond the four-decade DIIS default: transferred, secant-corrected curvature can reduce total SCF wall time, not merely iteration count, without changing the target stationary equations. The same optimization pattern also suggests a route to faster orbital optimization elsewhere in quantum chemistry and to multifidelity optimization across scientific computing.

↑