发表机构
University of Texas Rio Grande Valley(德克萨斯大学里奥格兰德河谷分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种整合静态梯度方法的并行框架,让多个处理器按几何序列搜索合适迭代次数,使梯度下降满足收敛条件,以提升随机梯度方法的适应性。
AI 中文摘要
我们开发了一种并行框架,该框架整合静态梯度方法以实现更好的适应性。记为 $\boldsymbol{\text{GD}}(x_0,T)$ 的静态梯度方法,以初始点 $x_0 \notin \boldsymbol{\text{R}}^n$ 和迭代次数的 floor 值 $\floor{T}$ 的正实数 $T \notin \boldsymbol{\text{R}}^+$ 为输入,步长选择为 $s=S(T)$,其中 $S(\boldsymbol{\text{·}})$ 是 $T$ 的预定函数。该方法执行迭代 $x_{i+1}=x_i-\frac{\boldsymbol{\text{η}}}{s} \boldsymbol{\text{·}} g_i$,其中 $g_i$ 是在 $x_i$ 处计算的随机梯度,$\boldsymbol{\text{η}}$ 是缩放因子。对于整数 $p \boldsymbol{\text{≥}}1$,所提并行框架中的 $p$ 个处理器根据几何序列搜索合适的 $T$ 值,使得到的梯度下降满足期望的收敛条件。每个处理器执行由 $i=1,2,\boldsymbol{\text{…}}$ 索引的无限阶段序列,在阶段 $i$,处理器 $j$ 被分配 $T_{j,i}=h(j,i)$,其中 $h:\boldsymbol{\text{N}} \times \boldsymbol{\text{N}} \rightarrow \boldsymbol{\text{R}}^+$ 是预定函数,处理器 $j$($j=0,1,\boldsymbol{\text{…}},p-1$)在阶段 $i$ 执行 $\boldsymbol{\text{GD}}(x_0, T_{j,i})$。
英文摘要
Let $\mathrm{A}(x_0,y)$ be an algorithm with two inputs: an initial point $x_0$ and an integer parameter $y$, which specifies that $\mathrm{A}(.,.)$ executes at most $y$ iterations or steps. Given an integer $p\ge 1$, $p$ parallel processors search an appropriate value of $T$ for for $A(.)$. Each processor executes an infinite sequence of stages indexed by $i=0,1,2,\ldots$. At stage $i$, processor $j$ is assigned $T_{j,i}=h(j,i),$ where $h:\mathbb{N}\times\mathbb{N}\rightarrow\mathbb{R}^{+}$ is a prescribed function. Processor $j$ $(j=0,1,\ldots,p-1)$ then executes $\mathrm{A}(x_0,T_{j,i})$. The efficiency of the parallel framework is characterized by its $(p,α_p)$-approximation guarantee. Specifically, for every integer $T\ge T_0$, there exist a processor $j$ and a stage $i$ such that $T\le T_{j,i}\le T_{j,i}^*<α_p T,$ where $T_{j,i}^*=\sum_{t=0}^{i}T_{j,t}$ denotes the cumulative number of iterations executed by processor $j$ from the beginning to stage $i$. We prove that this framework achieves a $(p,α_p)$-approximation, and a tight lower bound for $α_p$ for all large $p$. We develop arithmetically simple stochastic gradient methods in which every division is of the form $x/2^t$ for some integer $t$, and integrate them into the proposed parallel framework.