裁剪随机梯度下降的精确收敛速度
Precise Convergence Speed of Clipped SGD
浏览论文内容
中文总结 AI 辅助
针对裁剪随机梯度下降,提出收紧的收敛分析,简化证明并扩展适用范围,强化收敛准则并降低最终损失。
中文摘要 AI 辅助
我们针对$(L_0, L_1)$-光滑函数上的裁剪梯度下降提出了收紧的收敛分析,并给出了定量常数。基于Koloskova等人(2023)的思想,我们重构了若干情形划分,以揭示由$\ell_2$-投影的基本性质导出的偏差控制的核心作用,从而简化了证明。我们还将有效性范围从$\eta \leq 1 / (9 \beta)$扩展到$\eta < 1 /\beta$,其中裁剪常数$c$满足$\beta = L_0 + c L_1$,这与更传统的光滑函数分析一致。我们将收敛准则从$\left( \min_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$强化为$\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$,并保持匹配的速度,同时将最终可达到的损失从$\mathcal{O}(\min(\sigma^2/c, \sigma))$降低到更精确的$6 \min(\sigma^2 /c, 3 \sigma)$。
英文摘要
We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental properties of $\ell_2$-projection, simplifying proofs. We also extend the domain of validity from $η\leq 1 / (9 β)$ to $η< 1 /β$ where $β= L_0 + c L_1$ for clipping constant $c$, which matches the more traditional analysis of smooth functions. We strengthen the convergence criterion from $\left( \min_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ to $\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right)$ with matching speed, and lower the final achievable loss from $\mathcal{O}(\min(σ^2/c, σ))$ to the more precise $6 \min(σ^2 /c, 3 σ)$.
发表机构
- LAMSADE, Université Paris-Dauphine, PSL Research University(巴黎多菲纳大学PSL研究大学LAMSADE)
机构由 AI 辅助整理,请以论文原文为准。