arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17483math.OCcs.LG

弥合同质与异质异步优化之间的差距出奇地困难

Bridging the Gap Between Homogeneous and Heterogeneous Asynchronous Optimization Is Surprisingly Difficult

  • Applied AI Institute(应用人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Alexander Tyurin

AI总结:

本研究证明在常用相似性假设下,异质异步优化的悲观时间复杂度无法改进,并提出强插值与局部PL条件组合,达到与同质设置匹配的复杂度界。

AI中文摘要:

现代大规模机器学习任务通常需要多个工作节点、设备、CPU或GPU并行且异步地计算随机梯度,以训练模型权重。理论结果通常区分两种设置:(i)同质设置,其中所有工作节点都能访问相同的数据分布;(ii)异质设置,其中每个工作节点处理不同的数据分布。已知的这两种设置下的最优时间复杂度揭示了显著差距,异质情况下的保证要悲观得多。在这项工作中,我们研究了是否可以在不同假设下克服这些悲观的最优时间复杂度。令人惊讶的是,我们证明在广泛使用的一阶和二阶相似性假设下,对于任何随机算法,改进在理论上都是不可能的。然后我们转向插值区间,并证明仅凭弱插值假设也是不够的。最后,我们引入了不可约假设的最小组合,即强插值条件和局部Polyak-Lojasiewicz条件,以推导出新的时间复杂度界,该界与同质设置中已知最佳结果中对工作节点计算时间的依赖相匹配,而无需要求相同的数据分布。

英文摘要:

Modern large-scale machine learning tasks often require multiple workers, devices, CPUs, or GPUs to compute stochastic gradients in parallel and asynchronously to train model weights. Theoretical results typically distinguish between two settings: (i) the homogeneous setting, where all workers have access to the same data distribution, and (ii) the heterogeneous setting, where each worker operates on different data distributions. Known optimal time complexities in these settings reveal a significant gap, with far more pessimistic guarantees in the heterogeneous case. In this work, we investigate whether these pessimistic optimal time complexities can be overcome under different assumptions. Surprisingly, we show that improvement is provably impossible under widely used first- and second-order similarity assumptions for any randomized algorithm. We then turn to the interpolation regime and demonstrate that the weak interpolation assumption alone is also insufficient. Finally, we introduce a minimal combination of irreducible assumptions, strong interpolation and the local Polyak-Lojasiewicz condition, to derive a new time complexity bound that matches the dependence on worker computation times in the best-known result in the homogeneous setting, without requiring identical data distributions.

↑