发表机构
Kyoto University; NII LLMC; RIKEN AIP(京都大学; 国立情报学研究所LLMC; 理化学研究所革新智能统合研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对联邦回归中多最优头部导致的聚合歧义,提出自然选择规则(返回最接近广播头部的解),证明其收敛到Bures-Wasserstein重心,并给出基于目标均值与协方差交换的一轮修正以逼近集中式最优。
AI 中文摘要
在联邦平均中,局部目标可以允许多个最优头部,使得聚合结果取决于客户端返回的头部。我们研究了具有私有骨干网络和共享线性头部的联邦多元回归中的这种歧义性,使用一种无约束特征模型(UFM),该模型将训练样本特征视为自由变量。我们引入了一种自然的选择规则:每个客户端返回与广播头部最接近的最优头部。我们证明,在头部上带有消失的近端惩罚的全局最小化实现了这一规则。当客户端的最优Gram矩阵和初始共享Gram矩阵是正定时,共享Gram矩阵遵循一个封闭的递归,并收敛到客户端最优Gram矩阵的唯一Bures-Wasserstein重心。即使有了这种对齐,极限通常也不同于集中式最优Gram矩阵。我们将这一差距分解为三个正半定项,这些项源于客户端目标均值的差异、协方差异质性以及对齐头部的平均。基于一次性交换目标均值和协方差的修正,在精确局部优化和相同选择规则下,可以在一轮内恢复集中式最优Gram矩阵。我们在UFM中数值验证了这些结果,并在五个表格和五个图像回归数据集上使用具有特征正则化和长局部训练的深度网络测试了其预测。在这些实验中,普通训练接近预测的重心,而弱的近端惩罚改善了端点一致性,并产生了紧密遵循预测Gram动力学的轨迹。该修正使最终的Gram矩阵接近集中式UFM预测。
英文摘要
In federated averaging, local objectives can admit multiple optimal heads, making the aggregate depend on which heads clients return. We study this ambiguity in federated multivariate regression with private backbones and a shared linear head, using an unconstrained feature model (UFM) that treats training-sample features as free variables. We introduce a natural selection rule: each client returns the optimal head closest to the broadcast head. We show that global minimization with a vanishing proximal penalty on the head realizes this rule. When the clients' optimal Gram matrices and the initial shared Gram matrix are positive definite, the shared Gram matrix follows a closed recursion and converges to the unique Bures-Wasserstein barycenter of the clients' optimal Gram matrices. Even with this alignment, the limit generally differs from the centralized optimal Gram matrix. We decompose this gap into three positive-semidefinite terms arising from differences in client target means, covariance heterogeneity, and averaging the aligned heads. A correction based on a one-time exchange of target means and covariances recovers the centralized optimal Gram matrix in one round under exact local optimization and the same selection rule. We verify these results numerically in the UFM and test its predictions on five tabular and five image regression datasets using deep networks with feature regularization and long local training. In these experiments, ordinary training approaches the predicted barycenter, while a weak proximal penalty improves endpoint agreement and yields trajectories that closely follow the predicted Gram dynamics. The correction moves the final Gram matrices close to the centralized UFM prediction.
Comments47 pages, 17 figures, 9 tables