arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通信高效、鲁棒且实用的分布式与联邦优化的理论基础

Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization

Grigory Malinovsky

arXiv 2608.06563首次发表:更新:

AI 中文总结

本论文针对分布式与联邦优化的七大瓶颈,提出ProxSkip等算法,建立通信高效、鲁棒优化的理论框架,经实验验证其有效性。

AI 中文摘要

机器学习与优化协同发展,实际需求催生新理论,而理论突破又推动新应用。现代大规模训练依赖经典优化原理,但分布式系统的约束要求重新审视这些基础。本论文聚焦理论与实践交叉的七大挑战,重点研究联邦学习与分布式优化的关键瓶颈。其一,我们提出ProxSkip算法并证明局部梯度步长可加速通信,为这一广泛使用的启发式方法提供理论基础。其二,我们开发Variance Reduced ProxSkip算法,消除随机局部更新的邻域误差,同时平衡通信与局部计算。其三,我们证明在部分客户端参与下,局部步长仍能保持通信加速效果。其四,我们证明在异构场景下,服务器端步长与无放回采样可提升收敛性。其五,针对Random Reshuffling方法,我们证明压缩梯度差异而非梯度本身可获得更优的理论与实际性能。其六,我们确立通过梯度差异裁剪可同时实现拜占庭鲁棒性与部分客户端参与。最后,我们开发首个基于随机非对称链的低秩适应理论框架,为大模型微调提供新见解。在所有贡献中,我们引入新颖算法框架,在现实假设下建立严格保证,并通过数值实验为理论提供支撑。

英文摘要

Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. Modern large-scale training relies on classical optimization principles, but the constraints of distributed systems require these foundations to be reconsidered. This thesis addresses seven challenges at the intersection of theory and practice, focusing on key bottlenecks in federated learning and distributed optimization. First, we introduce ProxSkip and prove that local gradient steps can accelerate communication, providing a theoretical foundation for this widely used heuristic. Second, we develop Variance Reduced ProxSkip, which eliminates the neighborhood error of stochastic local updates while balancing communication and local computation. Third, we show that local steps retain their communication acceleration under partial client participation. Fourth, we prove that server-side stepsizes and sampling without replacement improve convergence in heterogeneous settings. Fifth, for Random Reshuffling, we demonstrate that compressing gradient differences rather than gradients yields better theoretical and practical performance. Sixth, we establish that Byzantine robustness and partial participation can be achieved simultaneously using gradient-difference clipping. Finally, we develop the first theoretical framework for low-rank adaptation based on randomized asymmetric chains, providing new insights into fine-tuning large models. Across these contributions, we introduce novel algorithmic frameworks, establish sharp guarantees under realistic assumptions, and support the theory with numerical experiments.

Commentsthesis, 7 chapters

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑