训练神经网络中的子空间列文伯格-马夸尔特算法
Subspace Levenberg Marquardt Algorithms in Training Neural Networks
浏览论文内容
中文总结 AI 辅助
本研究针对中小规模神经网络训练中LM算法开销随参数增长的问题,评估子空间LM算法在回归分类任务的表现,并与经典LM、SGD、Adam等方法对比性能。
中文摘要 AI 辅助
列文伯格-马夸尔特(LM)算法是一种知名的二阶方法,在训练中小规模神经网络(NN)时具备快速收敛性与强鲁棒性。然而,随着神经网络参数数量增长,其计算与内存开销会显著增加。为解决这一局限,研究人员提出了子空间方法,如克里洛夫子空间LM(KSLM)与混合子空间LM(HSLM),使二阶算法更高效。本研究评估了子空间LM算法在神经网络回归与分类任务中的表现,将子空间LM变体的性能与经典LM方法,以及随机梯度下降(SGD)、Adam等流行一阶算法进行对比。
英文摘要
The Levenberg-Marquardt (LM) algorithm is a well-known second-order method for rapid convergence and strong robustness when training small- to medium-sized neural networks (NNs). However, its computational and memory costs increase significantly as the number of parameters in an NN grows. To address this limitation, subspace methods have been proposed, such as the Krylov subspace LM (KSLM) and the hybrid subspace LM (HSLM), making second-order algorithms more efficient. In this work, we evaluate the subspace Levenberg-Marquardt algorithms for regression and classification tasks in neural networks. We compare the performance of subspace LM variants with the classical LM method, as well as other popular first-order algorithms, such as stochastic gradient descent (SGD) and Adam.
发表机构
- University of California, Davis(加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。