发表机构
School of Electrical, Computer and Energy Engineering, Arizona State University; Department of Electrical Engineering (ISY), Linköping University(亚利桑那州立大学电气、计算机与能源工程学院; 林雪平大学电气工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对流数据下带通信约束的分散优化问题,提出时间加权的结构化时变公式,分析分散一阶方法的跟踪误差界,通过数值实验验证了时间加权规则等因素对性能的影响。
AI 中文摘要
优化理论是智能决策的广泛应用工具。经典优化处理固定、时不变的目标函数,而许多现代应用处于动态环境中,数据按顺序到达,学习目标随时间演变,且常受分散数据和通信约束。受这些趋势驱动,我们通过结构化时变公式研究流数据下的分散优化,其中全局目标是网络上观测到的损失的时间加权平均。我们分析多迭代分散一阶方法,包括分散梯度下降。对于强凸且光滑的损失,我们通过收缩映射视角推导欧几里得范数跟踪误差的保证,所得界将跟踪误差分解为不动点跟踪分量和由分散化及数据异质性引起的偏差项。我们将分析专门化到均匀权重、指数贴现权重及其有限记忆窗口对应物。这些界明确刻画了时间加权规则、每步迭代预算、步长和网络连通性的作用。均匀加权产生阶为O(1/t)的消失不动点跟踪贡献,而贴现和窗口策略通常分别由贴现因子和有效记忆控制产生非消失跟踪底。在所有情况下,恒定步长下分散化会引起额外非零偏差底。数值实验验证了预测趋势。
英文摘要
Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant objective functions, many modern applications operate in dynamic environments where data arrive sequentially, and the learning objective evolves over time, often under decentralized data and communication constraints. Motivated by these trends, we study decentralized optimization from streaming data through a structured time-varying formulation in which the global objective is a temporally weighted average of losses observed across the network. We analyze multi-iteration decentralized first-order methods, including decentralized gradient descent. For strongly convex and smooth losses, we develop guarantees for the Euclidean-norm \emph{tracking error} through a contraction-mapping viewpoint. The resulting bounds decompose the tracking error into a fixed-point tracking component and a bias term induced by decentralization and data heterogeneity. We specialize our analysis to uniform and exponentially discounted weights, as well as their finite-memory \emph{windowed} counterparts. The bounds explicitly characterize the roles of the temporal weighting rule, per-step iteration budget, step size, and network connectivity. Uniform weighting yields a vanishing fixed-point tracking contribution of order $\mathcal O(1/t)$, whereas discounted and windowed strategies generally induce non-vanishing tracking floors governed by the discount factor and effective memory, respectively. In all cases, decentralization induces an additional non-zero bias floor under a constant step size. Numerical experiments illustrate the predicted trends.