发表机构
Università di Pisa; Université Lumière Lyon 2(比萨大学; 里昂第二大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从遍历理论视角研究一大类随机优化算法的渐近性质,证明随机梯度下降等算法达到极小值邻域的命中时间围绕均值呈指数分布,该均值由目标平稳测度倒数决定。
AI 中文摘要
机器学习,尤其是深度学习,涉及求解大规模非凸优化问题。文献中已提出多种算法,这些算法在困难实例上似乎能达到令人满意的实际效率,其中随机梯度法最为基础,但在许多学习任务上仍优于更新的算法。关于当前深度学习方法的一个主要开放问题是理解其收敛性质。沿着先前关于梯度类算法长时间行为的研究思路,特别是Azizian等人最近的贡献,我们提出了一种新的方法,从遍历理论的角度研究一大类方法的渐近性质。我们的主要结果包括对随机优化算法达到某个极小值点小邻域的期望时间的研究,并表明该到达时间围绕其平均值呈指数分布,该平均值由目标的平稳测度的倒数给出。对随机梯度噪声的假设包括高斯假设和次指数假设。
英文摘要
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most rudimentary, while still outperforming more recent algorithms at a number of learning tasks. A major open question about the current methods used in deep learning is to understand their convergence properties. Following a line of previous works about the long time behavior of gradient-type algorithms, %and in particular the recent contributions from Azizian et al., we present a new approach for studying the asymptotic properties of a wide family of methods from an ergodic theoretical viewpoint. Our main results include a study of the expected time for a stochastic optimisation algorithm to reach a certain small neighborhood of a minimizer and show that this reaching time distributes exponentially around its average, which is given by the inverse of the stationary measure of the target. The assumptions on the Stochastic Gradient noise include the Gaussian and the Sub-Exponential assumptions.
Comments24 pages