arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式随机优化中通过树路由实现的最优网络依赖性

Optimal Network Dependence in Distributed Stochastic Optimization via Tree Routing

Runze You, Shi Pu

arXiv 2609.18101首次发表:更新:

AI 中文总结

本文通过图直径与逆谱间隙的关系,提出树路由梯度跟踪方法,在分布式随机非凸优化中实现直径最优的网络依赖性,无需预定混合矩阵。

AI 中文摘要

通信是分布式优化的核心瓶颈,但其影响通常通过所选混合矩阵的谱间隙来概括。由于该间隙既取决于链路权重又取决于拓扑结构,因此可能掩盖网络的固有影响。我们通过将图直径与逆谱间隙相关联来阐明这一区别。我们建立了一个普适的单侧界,该界在常数因子内是紧的,而跨多个图族的构造表明,不存在仅基于拓扑的逆命题或普适的双侧缩放律。这促使我们采用基于类别的网络最优性观点:直径相关的保证可以在图类上统一评估,而谱间隙最优性则需要同时指定图-矩阵对类别和可容许算法类别。基于由此产生的基于直径的极小极大基准,我们引入了树路由梯度跟踪(Tree-RGT),一种用于随机非凸优化的分布式方法。该方法在根最短路径生成树上流水线化模型传播和梯度聚合。每次迭代使用一轮单跳通信和每个智能体一次随机梯度评估,无需内部共识或多重八卦。在光滑性和无偏有界方差预言机下,Tree-RGT在具有指定大小和直径的连通图上实现了均匀最优的网络依赖性。在$\mathcal{O}(nD_{\mathcal{G}}^2)$次瞬态迭代后,它恢复了集中式随机缩放,其中$n$是智能体数量,$D_{\mathcal{G}}$是图直径。这种依赖于网络的瞬态迭代界改进或匹配了代表性分布式方法的报告结果。因此,拓扑感知路由在不依赖预定混合矩阵的情况下实现了直径最优保证。

英文摘要

Communication is a central bottleneck in distributed optimization, but its effect is often summarized by the spectral gap of a chosen mixing matrix. Since this gap depends on link weights as well as topology, it can obscure the intrinsic effect of the network. We clarify this distinction by relating graph diameter to the inverse spectral gap. We establish a universal one-sided bound that is tight up to constant factors, while constructions across several graph families show that no topology-only converse or universal two-sided scaling law exists. This motivates a class-based view of network optimality: diameter-dependent guarantees can be assessed uniformly over graph classes, whereas spectral-gap optimality requires both the graph--matrix pair class and the admissible algorithm class to be specified. Based on the resulting diameter-based minimax benchmark, we introduce Tree-Routed Gradient Tracking (Tree-RGT), a distributed method for stochastic nonconvex optimization. The method pipelines model dissemination and gradient aggregation over a rooted shortest-path spanning tree. Each iteration uses one round of one-hop communication and one stochastic-gradient evaluation per agent, without inner consensus or multi-gossip. Under smoothness and unbiased bounded-variance oracles, Tree-RGT achieves uniformly optimal network dependence over connected graphs of prescribed size and diameter. It recovers centralized stochastic scaling after $\mathcal{O}(nD_{\mathcal{G}}^2)$ transient iterations, where $n$ is the number of agents and $D_{\mathcal{G}}$ is the graph diameter. This network-dependent transient-iteration bound improves upon or matches those reported for representative distributed methods. Thus, topology-aware routing attains diameter-optimal guarantees without relying on a prescribed mixing matrix.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑