发表机构
Tianjin University(天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FedMIX-P通过每步混合共享与局部预条件子,降低联邦训练中的算子不匹配,理论保证O(R^{-1/2})平稳性,实验在视觉与语言任务上提升准确率最高19.47个百分点。
AI 中文摘要
自适应预条件子能够加速模型训练,但即使梯度在同一模型上评估,异构客户端几何特性也可能使联邦更新产生偏差。仅靠轮次开始时的同步无法防止这种不匹配在本地训练过程中重新出现。我们提出FedMIX-P,该方法在每一步本地更新中混合共享预条件子和局部预条件子,在保留本地适应性的同时,将均方算子不匹配度降低λ²倍。针对具有随机梯度和部分参与的光滑非凸目标,我们利用合适的步长和依赖于训练视野的混合权重,建立了O(R^{-1/2})的平稳性界,且无需要求各局部预条件子相互收敛。一个双客户端反例表明,固定的正混合权重可能保留非平稳不动点。该理论涵盖有界线性对称正定预条件子。在视觉和语言任务上使用SOAP、Sophia和Muon变体的实验表明,与相应的本地优化器相比,所提方法有改进,包括在60M至350M语言模型上准确率最高提升19.47个百分点,并降低验证损失。完整的非线性和基于动量的更新需要单独分析。
英文摘要
Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \texttt{FedMIX-P}, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of $λ^2$. For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an $O(R^{-1/2})$ stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to $19.47$ percentage points and lower validation loss for 60M--350M language models. Full nonlinear and momentum-based updates require separate analysis.