发表机构
University of California, Davis(加利福尼亚大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出带充足性监测器的自适应混合子空间LM算法,通过构建低维子空间解耦步长接受与阻尼调整,在保证收敛性的同时降低大规模最小二乘问题的迭代计算成本,在神经网络训练中表现优异。
AI 中文摘要
列文伯格-马夸尔特(LM)算法是求解非线性最小二乘问题最广泛使用的方法,它结合了最速下降法的鲁棒性与高斯-牛顿法的快速局部收敛性。然而,对于大规模问题,其计算成本可能过高,因为每次迭代都需要求解一个大型阻尼线性系统,且传统的步长接受策略可能会因阻尼参数调整而需要重复求解。尽管存在这一计算挑战,许多大规模最小二乘问题呈现出有效的低维结构,仅存在少量受数据强烈影响的参数空间方向。我们提出一种自适应混合子空间列文伯格-马夸尔特(HSLM)算法,该算法从梯度、记忆、克里洛夫子空间及随机曲率信息的互补来源构建低维子空间,并在该子空间内计算谱阻尼LM步。该方法的一个显著特征是确定性充足性监测器,它量化降维空间捕获的下降信息,并在必要时自适应丰富子空间。步长接受与阻尼调整解耦:Armijo回溯法确定接受的步长,而实际与预测下降的比值仅用于更新阻尼参数,从而避免了步长接受过程中重复求解阻尼系统。对于HSLM算法,我们建立了全局收敛到平稳性的结论,并证明其局部线性和超线性收敛性。在神经网络训练问题上的数值实验表明,HSLM的收敛行为可与经典及克里洛夫子空间LM(KSLM)相媲美,同时大幅降低了每次迭代的计算成本,且随着参数维度的增长,优势愈发明显。
英文摘要
The Levenberg-Marquardt (LM) algorithm is the most widely used method for solving nonlinear least-squares problems, as it combines the robustness of steepest descent with the fast local convergence of the Gauss-Newton method. However, its computational cost can become prohibitive for large-scale problems because each iteration requires solving a large damped linear system, and conventional step acceptance strategies may require repeated solves as the damping parameter is adjusted. Despite this computational challenge, many large-scale least-squares problems exhibit effective low-dimensional structure, with only a small number of parameter-space directions strongly informed by the data. We propose an adaptive hybrid subspace Levenberg-Marquardt (HSLM) algorithm that constructs a low-dimensional subspace from complementary sources of gradient, memory, Krylov-subspace, and randomized curvature information and computes a spectrally damped LM step within this subspace. A distinguishing feature of the method is a deterministic adequacy monitor that quantifies how much descent information is captured by the reduced space and adaptively enriches the subspace when necessary. Step acceptance is decoupled from damping adjustment: Armijo backtracking determines the accepted step length, while the ratio of actual to predicted reduction is used solely to update the damping parameter, thereby avoiding repeated damped-system solves during step acceptance. For the HSLM algorithm, we establish global convergence to stationarity and prove local linear and superlinear convergence. Numerical experiments on neural-network training problems show that HSLM achieves convergence behavior comparable to classical and Krylov subspace LM (KSLM) while substantially reducing per-iteration computational cost, with increasing advantages observed as the parameter dimension grows.
Comments28 pages, 5 figures