发表机构
UFSC Blumenau; School of Applied Mathematics, FGV; Instituto Tecnológico Autónomo de México, ITAM(布卢梅瑙联邦圣卡塔琳娜大学; 巴西基金会应用数学学院; 墨西哥自治理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对凸二次极小化,提出带松弛的 RelaxME 和带动量的 MomME 两种 ME 变体,证明其线性收敛性,数值实验显示 MomME 迭代次数少、CPU 时间短,为 ME 拓展应用提供可能。
AI 中文摘要
椭心方法(ME)是一种新近提出的用于无约束极小化的技术,其迭代仅依赖一阶信息,核心是构建合适的二维椭圆以捕捉问题的固有病态性,将椭圆中心作为下一个迭代点,这也是该方法名称的由来。已证明当目标函数为光滑强凸函数时,ME 具有线性收敛性;目标函数为二次函数的特殊情形在 ME 的首篇论文中被研究,也是本文的研究主题。本文提出 ME 的两个变体:在每次迭代中添加额外松弛步骤的 RelaxME,以及嵌入动量的 MomME。我们证明了 RelaxME 和 MomME 的线性收敛性,特别地,当二次型矩阵仅有两个不同特征值时,RelaxME、MomME(以及 ME)在一次迭代内即可收敛。最后,我们给出数值实验结果,在五个优化问题(及定义这些问题的不同参数组合)上,将 ME、RelaxME、MomME 与另外三种优化器——共轭梯度法[10]、长步长 Barzilai-Borwein 梯度法[2]、自适应谱步长梯度法[7]进行对比。在多数算例中,MomME 的迭代次数最少,共轭梯度法与 MomME 的 CPU 时间最少且二者相近,这为带动量的 ME 在更广泛场景中的应用提供了乐观可能。
英文摘要
The method of ellipcenters (ME) is a recent technique developed for unconstrained minimization. Its iteration relies only on first order information and consists of building a suitable two-dimensional ellipse that tries to capture intrinsic ill-conditioning of the problem. The center of the ellipse is taken as the next iterate, which justifies the name of the method. ME was already shown to converge linearly when the objective function is smooth and strongly convex. The special case when the objective is quadratic was studied in the first paper on ME and is also the subject of our work here. In this paper, we propose two variants of ME: RelaxME which adds in ME an additional relaxation step at each iteration and MomME which embeds momentum in ME. We prove linear convergence of RelaxME and MomME and, in particular, we show convergence of RelaxME and MomME (and also of ME) in one iteration when the matrix of the quadratic form has only two distinct eigenvalues. Finally, we provide the results of numerical experiments which compare on five optimization problems (and different combinations of parameters defining these problems) ME, RelaxME, and MomME, with 3 other optimizers: the conjugate gradient method [10], Barzilai and Borwein gradient method with long step [2], and the gradient method with adaptive spectral step length [7]. On most instances, MomME provides the smallest number of iterations and conjugate gradient and MomME provide the smallest CPU times and similar CPU times. This opens optimistic possibilities for ME with momentum in broader settings.