FedeRage:一般客户端漂移下可证明收敛的不可知联邦学习
FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift
浏览论文内容
中文总结 AI 辅助
针对客户端参与概率未知且数据异构的联邦学习,提出风险规避平均算法FedeRage,嵌入CVaR提升高损失和低频客户端权重,实现可证明收敛并在准确率、公平性和速度上超越现有方法。
中文摘要 AI 辅助
联邦学习(FL)能够在无需共享原始数据的情况下进行协作模型训练,但在非独立同分布(non-IID)数据和随机客户端参与的情况下,其性能会下降。基于经典联邦平均(FedAvg)的改进方法通常预设服务器已知客户端参与概率,而这在部署系统中很少成立。我们首先讨论并刻画了当参与完全未知、可能高度偏斜且各轮规模可变时,\textit{分布不可知}的FedAvg实际求解的优化问题:统一聚合被证明是在一个由参与诱导的边际加权的明确定义的随机目标上,以凸且可能非光滑损失的标准$\mathcal{O}(1/\sqrt{T})$速率进行最小化。基于这一刻画,我们提出了\textit{联邦风险规避平均}(\textsc{FedeRage}),这是FedAvg的一种风险规避扩展,在自然的分布鲁棒优化(DRO)框架内将\textit{条件风险价值}(CVaR)嵌入局部目标。\textsc{FedeRage}隐式地提高了高损失和低频参与客户端的权重,同时仅为每个客户端添加\textit{单个标量},并实现了$\mathcal{O}(\kappa/\sqrt{T})$的速率,其中因子$\kappa$是风险规避“代价”的上界。与基于最优传输的聚合对齐方案(需要可用分布作为输入)相比,\textsc{FedeRage}对其保持不可知。在三个异构基准上的多项实验表明,在准确性、公平性和收敛速度方面均优于现有最先进方法。
英文摘要
Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. Remedies built on classical Federated Averaging (FedAvg) typically presuppose that client participation probabilities are known to the server, which is rarely the case in deployed systems. We first discuss and then characterize the optimization problem that \emph{distributionally agnostic} FedAvg actually solves when participation is entirely unknown, possibly highly skewed, and of variable size across rounds: uniform aggregation is shown to minimize a well-defined stochastic objective, weighted by the participation-induced marginal, at a standard $\mathcal{O}(1/\sqrt{T})$ rate for convex and possibly nonsmooth losses. Building on this characterization, we propose \emph{Federated Risk-Averse Averaging} (\textsc{FedeRage}), a risk-averse extension of FedAvg that embeds the \emph{Conditional Value-at-Risk} (CVaR) into the local objective within a natural distributionally robust optimization (DRO) framework. \textsc{FedeRage} implicitly upweights high-loss and infrequently participating clients while adding only a \emph{single scalar per-client}, and admits an $\mathcal{O}(κ/\sqrt{T})$ rate in which the factor $κ$ is the upper bound on the ``price" of risk aversion. In contrast with aggregation-alignment schemes based on optimal transport, which require the availability distribution as an input, \textsc{FedeRage} remains agnostic to it. Several experiments on three heterogeneous benchmarks indicate consistent improvements over state-of-the-art methods in accuracy, fairness, and convergence speed.