基于记忆持久性的随机子空间梯度下降
Gradient Descent with Stochastic Subspaces via Persistence of Memory
- National University of Singapore(新加坡国立大学)
- RIKEN(理化学研究所)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出“记忆持久性”技术改进随机子空间梯度下降,利用与梯度弱相关的引导向量指导子空间生成,可长时间固定并稀疏或小批量场景下廉价获取,首次给出稀疏函数理论分析,并展示梯度估计与Hessian低位特征向量对齐,实现一次性计算引导向量,提升计算效率。
中文摘要 AI 辅助
随机子空间方法作为基于梯度下降的大规模优化问题技术,尤其在分布式设置中,已获得广泛关注。本文中,我们引入了“记忆持久性”技术,以大幅扩展和改进随机子空间方法。为此,我们利用一个与梯度弱相关的向量,为随机子空间的生成过程提供引导结构,沿该子空间进行下降。该引导向量可在大量迭代中保持固定,仅在宽间隔后刷新(关于间隔大小,我们可根据问题参数提供保证)。在重要的机器学习设置中,例如体现稀疏性或小批量结构的优化问题,我们展示了通过利用问题的结构化属性,可以以有效且计算廉价的方式获得引导向量。在此过程中,我们建立了据我们所知对经典SSD方法在稀疏函数上的首次理论分析。在最优点的局部邻域内,我们展示了我们的梯度估计与Hessian的低位特征向量之间的对齐现象,允许一次性计算引导向量,这使得该方法即使在非结构化目标场景中也具有计算优势。
英文摘要
Stochastic subspace methods have gained popularity as gradient descent based techniques for large scale optimisation problems, especially in distributed settings. In this paper, we introduce the technique of "persistence of memory" to greatly extend and improve the random subspace methods. To this end, we leverage a vector that is only weakly correlated with the gradient in order to provide a guiding structure to the generative process of the random subspace along which the descent is going to take place. This guidance vector may be fixed for a large number of iterations, only to be refreshed at wide intervals (on whose size we can provide guarantees in terms of problem parameters). In important machine learning settings, such as optimisation problems embodying sparsity or a minibatch structure, we show that the guidance vector can be obtained in an effective and computationally inexpensive manner by leveraging the structured properties of the problem. En route, we establish to our knowledge the first theoretical analysis of classical SSD methods for sparse functions. In a local neighbourhood of the optimum, we demonstrate an alignment phenomenon of our gradient estimates with a low-lying eigenvector of the Hessian, allowing a once-for-all computation of the guidance vector which renders the method computationally favourable even in scenarios with unstructured objectives.