arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于RNN训练中非凸非光滑多复合优化的序贯正则化分段仿射算法

A sequential regularized piecewise affine algorithm for nonconvex nonsmooth multicomposite optimization in RNN training

Lingzi Jin, Xiao Wang, Xiaojun Chen

arXiv 2609.18325首次发表:更新:

发表机构

Hong Kong Polytechnic University; Sun Yat-sen University(香港理工大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出序贯正则化分段仿射算法SRPA,用于训练RNN的非凸非光滑多复合优化,保证收敛到二阶d-稳定点,复杂度为O(ε^{-2}),实验验证其性能优于现有算法。

AI 中文摘要

本文关注一类用于训练RNN(循环神经网络)的非凸非光滑多复合优化问题。我们首先建立了易于验证的条件,在这些条件下,问题的近似一阶d-稳定点保证是近似二阶d-稳定点。随后,我们提出了一种序贯正则化分段仿射算法SRPA,该算法在每次迭代中最小化一个正则化的分段仿射函数,其中融入了主动集策略的思想以降低每次迭代的计算成本。利用上述二阶d-稳定性的条件,我们建立了SRPA全局收敛到二阶d-稳定点的结论,以及获得ε-近似二阶d-稳定点的复杂度界为O(ε^{-2})。最后,在合成和真实世界数据集上训练RNN的数值实验表明,与最先进的算法相比,SRPA表现出有前景的性能。

英文摘要

This paper focuses on a class of nonconvex nonsmooth multicomposite optimization problems for training RNNs (Recurrent Neural Networks). We first establish easily verifiable conditions under which an approximate first-order d-stationary point of the problem is guaranteed to be an approximate second-order d-stationary point. Subsequently, we propose a sequential regularized piecewise affine algorithm, SRPA, that minimizes a regularized piecewise affine function at each iteration, where the idea of active-set strategy is incorporated to reduce the computational cost per iteration. Leveraging the aforementioned conditions for second-order d-stationarity, we establish the global convergence of SRPA to second-order d-stationary points and the complexity bound of $ \mathcal{O}(ε^{-2}) $ for obtaining an $ ε$-approximate second-order d-stationary point. Finally, numerical experiments for training RNNs on synthetic and real-world datasets demonstrate the promising performance of SRPA compared with state-of-the-art algorithms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑