发表机构
Hong Kong Polytechnic University; Sun Yat-sen University(香港理工大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出序贯正则化分段仿射算法SRPA,用于训练RNN的非凸非光滑多复合优化,保证收敛到二阶d-稳定点,复杂度为O(ε^{-2}),实验验证其性能优于现有算法。
AI 中文摘要
本文关注一类用于训练RNN(循环神经网络)的非凸非光滑多复合优化问题。我们首先建立了易于验证的条件,在这些条件下,问题的近似一阶d-稳定点保证是近似二阶d-稳定点。随后,我们提出了一种序贯正则化分段仿射算法SRPA,该算法在每次迭代中最小化一个正则化的分段仿射函数,其中融入了主动集策略的思想以降低每次迭代的计算成本。利用上述二阶d-稳定性的条件,我们建立了SRPA全局收敛到二阶d-稳定点的结论,以及获得ε-近似二阶d-稳定点的复杂度界为O(ε^{-2})。最后,在合成和真实世界数据集上训练RNN的数值实验表明,与最先进的算法相比,SRPA表现出有前景的性能。
英文摘要
This paper focuses on a class of nonconvex nonsmooth multicomposite optimization problems for training RNNs (Recurrent Neural Networks). We first establish easily verifiable conditions under which an approximate first-order d-stationary point of the problem is guaranteed to be an approximate second-order d-stationary point. Subsequently, we propose a sequential regularized piecewise affine algorithm, SRPA, that minimizes a regularized piecewise affine function at each iteration, where the idea of active-set strategy is incorporated to reduce the computational cost per iteration. Leveraging the aforementioned conditions for second-order d-stationarity, we establish the global convergence of SRPA to second-order d-stationary points and the complexity bound of $ \mathcal{O}(ε^{-2}) $ for obtaining an $ ε$-approximate second-order d-stationary point. Finally, numerical experiments for training RNNs on synthetic and real-world datasets demonstrate the promising performance of SRPA compared with state-of-the-art algorithms.