arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07040math.OC

通过重参数化进行优化:具有奇异几何的时间扭曲镜像流

Optimization by Reparametrization: Time-Warped Mirror Flows with Singular Geometry

Cristian Vega, Cesare Molinari, Lorenzo Rosasco, Silvia Villa

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过将复杂模型视为简单模型的过参数化,研究其梯度流与原始参数上的镜像流对应关系,建立了适定性和收敛性,并揭示了过参数化的隐式正则化作用。

中文摘要 AI 辅助

解决数据驱动问题需要定义复杂模型并将其拟合到数据上,神经网络就是一个激励性的例子。拟合过程可以被视为一个优化问题,该问题通常是非凸的,因此难以推导出优化保证。一个机会在于将所关注的模型视为某个更简单模型的冗余重参数化——即过参数化——对于该更简单模型,优化结果更容易实现。在本文中,在形式化上述思想之后,我们重新审视了一些近期结果并推导出新的结果。特别地,我们考虑某些线性过参数化类别的梯度流,并表明它们对应于原始参数上的适当镜像流。我们的主要贡献涉及对后者的研究,我们为其建立了适定性和收敛性。提供了几个具体实例,包括全连接线性网络和具有权重归一化的网络。这些结果揭示了过参数化在优化中的作用以及相应的隐式正则化性质。

英文摘要

Solving data-driven problems requires defining complex models and fitting them to data, neural networks being a motivating example. The fitting procedure can be seen as an optimization problem, which is often non-convex, and hence optimization guarantees are hard to derive. An opportunity is provided by viewing the model of interest as a redundant reparameterization--an overparameterization--of some simpler model for which optimization results are easier to achieve. In this paper, after formalizing the above idea, we revisit some recent results and derive new ones. In particular, we consider the gradient flow of some classes of linear overparameterizations and show that they correspond to suitable mirror flows on the original parameters. Our main contribution relates to the study of the latter, for which we establish well-posedness and convergence. Several specific instances are provided, including fully connected linear networks and networks with weight normalization. The results yield insight into the role of overparameterization in optimization and corresponding implicit regularization properties.

发表机构

  • Universidad Técnica Federico Santa María(智利圣玛丽亚技术大学)
  • Università di Genova(热那亚大学)
  • Istituto Italiano di Tecnologia(意大利理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑