arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视基于符号的分布式方差缩减

Revisiting Distributed Sign-Based Variance Reduction

Wei Jiang, Zechao Li, Lijun Zhang

arXiv 2609.18656首次发表:更新:

发表机构

Nanjing University of Science and Technology; Nanjing University(南京理工大学; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对分布式符号方法在异构数据下的偏差问题,提出通过无偏压缩递归梯度增量跟踪全局梯度,实现了非凸随机和有限和优化的最优收敛速率,并匹配集中式样本复杂度。

AI 中文摘要

基于符号的方法在分布式环境中降低了通信成本,但当数据异构时,聚合局部符号可能会引入偏差。因此,现有的基于符号的方差缩减方法无法获得最优收敛速率。在本文中,我们解决了这一问题,并为非凸随机优化和有限和优化都获得了最优速率。我们首先给出一个反例,表明即使使用精确的局部梯度,多数投票也可能无法接近驻点。受此限制的启发,我们提出通过无偏压缩递归梯度增量来在服务器端跟踪全局梯度。因此,我们获得了关于ℓ1范数的收敛速率O(√(d/K)+√d(a/(nK))^{1/3})和关于ℓ2范数的O(√(a/K)+√a/(nK)^{1/3})。这里,K是迭代次数,n是工作节点数,d是维度,a=1+ω,其中ω表示压缩器的相对方差。对于具有M个分量的有限和问题,我们将周期性精确梯度刷新与压缩的分量梯度差异相结合。对于ℓ1和ℓ2梯度范数至多ε的情况,所得的总样本复杂度分别为O(M+d√(aM)ε^{-2})和O(M+a√M ε^{-2}),与集中式设置中的相应界相匹配。

英文摘要

Sign-based methods reduce communication costs in distributed environments, but aggregating local signs can introduce bias when data are heterogeneous. As a result, existing sign-based variance reduction methods fail to obtain the optimal convergence rates. In this paper, we solve this problem and obtain optimal rates for both nonconvex stochastic and finite-sum optimization. We first give a counterexample showing that majority voting can fail to approach stationary points even with exact local gradients. Motivated by this limitation, we propose tracking the global gradient at the server through unbiased compression of recursive gradient increments. As a result, we can obtain the convergence rates of $O(\sqrt{d/K}+\sqrt d (a/(nK))^{1/3})$ for the $\ell_1$-norm and $O(\sqrt{a/K}+\sqrt a/(nK)^{1/3})$ for the $\ell_2$-norm. Here, $K$ is the iteration number, $n$ is the number of workers, $d$ is the dimension, and $a=1+ω$, with $ω$ denoting the compressor's relative variance. For finite-sum problems with $M$ components, we combine periodic exact gradient refreshes with compressed component-gradient differences. The resulting total sample complexities are $O(M+d\sqrt{aM}ε^{-2})$ and $O(M+a\sqrt M\ epsilon^{-2})$ for $\ell_1$ and $\ell_2$ gradient norms at most $ε$, matching the corresponding bounds in centralized settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑