arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CTP-FL:联邦学习的公共轨迹梯度预测

CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning

Junkang Liu

arXiv 2609.35130首次发表:更新:

发表机构

Tianjin University(天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CTP-FL通过让所有客户端沿共享预测路径评估全局梯度,实现通信高效的联邦学习,无需有界梯度假设,并提供了非凸目标下的收敛界。

AI 中文摘要

通信高效的联邦优化通常在服务器更新之间花费多次梯度评估。现有的局部更新方法利用这些计算在每个客户端上推进一个独立的模型。然而,在异构数据下,这些模型在不同位置评估梯度,使得聚合更新难以解释为全局目标的梯度。我们研究相同计算预算的另一种用途:\u201c沿共享的预测路径评估全局目标\u201d。我们提出公共轨迹预测联邦学习(CTP-FL)。在每一轮中,所有客户端从当前全局模型和之前的聚合方向构造相同的查询点序列,沿该序列评估K个随机梯度,并上传它们的平均值。然后服务器执行一次全局更新。因此,CTP-FL在每个客户端使用K个小批量梯度,并在每个通信方向上传输一个模型大小的向量,与全参与FedAvg-M的每轮计算和通信量相匹配。共享查询点使得聚合方向成为沿预测路径的平均\u201c全局\u201d梯度的无偏估计量。与当前模型处梯度的剩余偏差由路径长度控制,无需假设客户端梯度差异有界或梯度有界。对于光滑非凸目标,我们在全参与下建立了$\u039f(\u221a(L\u0394\u03c3²/(NKR))+L\u0394/R)$的平均平稳性界。该分析隔离了一个可检验的权衡:延长预测路径提供更多前瞻性梯度信息,但增加了其位移偏差。

英文摘要

Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these models evaluate gradients at different locations, making the aggregated update difficult to interpret as a gradient of the global objective. We study an alternative use of the same computation budget: \emph{evaluate the global objective along a shared, predicted path}. We propose Common-Trajectory Predictive Federated Learning (\texttt{CTP-FL}). At each round, all clients construct the same sequence of query points from the current global model and the previous aggregated direction, evaluate $K$ stochastic gradients along this sequence, and upload their average. The server then performs a single global update. Thus, \texttt{CTP-FL} uses $K$ mini-batch gradients per client and one model-sized vector in each communication direction, matching the per-round computation and communication of full-participation FedAvg-M. Shared query points make the aggregated direction an unbiased estimator of the average \emph{global} gradient along the predicted path. The remaining discrepancy from the gradient at the current model is controlled by the path length, without assuming bounded client-gradient dissimilarity or bounded gradients. For smooth non-convex objectives, we establish an $\mathcal{O}\!\left( \sqrt{LΔσ^2/(NKR)}+LΔ/R \right)$ average-stationarity bound under full participation. The analysis isolates a testable trade-off: extending the prediction path provides more forward-looking gradient information but increases its displacement bias.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑