发表机构
ESCP Business School(ESCP商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文分析带初始正则化的随机梯度下降(SGDIR),推导其平方损失下的无维数上界,建立匹配的下界,通过实例级比较证明其期望超额风险不大于岭回归,数值实验验证理论结果。
AI 中文摘要
我们分析了一种带初始正则化的随机梯度下降(SGDIR)变体,并推导了其平方损失下期望超额风险的无维数上界。在无噪声情况下,我们在矩、源和容量假设下,为平均SGDIR和非平均SGDIR都得到了新的界。对于源参数的特定取值,这些界的阶为$m^{-2}\log^{2}m$,其中训练样本数的阶为$m$;对于源参数的另一取值,当容量参数超过$\epsilon^{-1}$时,对任意$\epsilon>0$,我们得到了阶为$m^{-3+\epsilon}$的界。我们还建立了一个下界,在某些区域内与我们的上界匹配,仅差一个多对数因子。在有噪声情况下,我们提供了SGDIR与岭回归的实例级比较。在一般假设和正则化参数的适度下界下,我们证明SGDIR的期望超额风险不大于岭回归的期望超额风险,仅差一个多对数因子。在合成数据和真实数据上的数值实验与我们的理论结果一致。
英文摘要
We analyze a variant of stochastic gradient descent with initial regularization (SGDIR) and derive dimension-free upper bounds on its expected excess risk for the squared loss. In the noiseless case, we obtain new bounds for both averaged and non-averaged SGDIR under moment, source, and capacity assumptions. For a particular value of the source parameter, these bounds are of order $m^{-2}\log^{2}m$, where the number of training samples is of order $m$. For another value of the source parameter, we obtain, for any $ε>0$, bounds of order $m^{-3+ε}$, provided that the capacity parameter exceeds $ε^{-1}$. We also establish a lower bound that matches our upper bounds in certain regimes up to a polylogarithmic factor. In the noisy case, we provide an instance-based comparison between SGDIR and ridge regression. Under general assumptions and a mild lower bound on the regularization parameter, we show that the expected excess risk of SGDIR is no larger than that of ridge regression, up to a polylogarithmic factor. Numerical experiments on synthetic and real data are consistent with our theoretical findings.
Comments33 pages