arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

向群体提出正确的问题:用于联邦学习的偏差消除权重

Asking the Crowd the Right Question: Bias-Cancelling Weights for Federated Learning

Ilya Kuruzov, Dmitrii Vishovan, Kirill Novoselov, Yuriy Dorn, Darina Dvinskikh, Alexander Gasnikov

arXiv 2610.04671首次发表:更新:

发表机构

Innopolis University; Moscow Independent Research Institute of Artificial Intelligence; Lomonosov Moscow State University; Higher School of Economics; Trusted AI Research Center, RAS; Steklov Mathematical Institute, RAS(Innopolis大学; 莫斯科人工智能独立研究院; 莫斯科罗蒙诺索夫国立大学; 高等经济学院; 俄罗斯科学院可信人工智能研究中心; 俄罗斯科学院斯捷克洛夫数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CROWD算法,将联邦学习中的固定权重视为群体智慧工具,通过从优化轨迹中读取客户端分歧来估计偏差并消除,从而在无需额外成本下达到贝叶斯极小极大最优,并在真实数据上优于按噪声方差加权。

AI 中文摘要

联邦目标函数是客户端风险的加权和,而这些权重几乎总是预先固定的。我们转而将它们视为群体智慧机制的唯一工具:客户端是对同一真相的有噪声的观察,每个客户端通过一个在群体中无偏的独立失真来观察真相。最优权重与客户端的误差能量成反比,这是经典结论;我们从该答案所预设的问题出发:哪些能量属于那里,以及群体能否从自身中恢复这些能量。在真相上的超额风险恰好与总体偏差能量的阶相同,因此任何优化器都无法修复一个糟糕的权重向量;真相本身仅可识别到线性倾斜,因此减去估计的客户端偏差可证明地恢复均匀加权。然而,期望的每客户端二阶矩可以从群体分歧的规律中通过一个良态线性反演精确识别,这是随机效应元分析无法采取的步骤,因为一个来源只报告一次;其实现对应量可估计到算法可测量的非相干性下限。这产生了CROWD,它从优化轨迹中读取分歧,无需额外成本,并在相同常数下匹配贝叶斯极小极大下界:随着时间范围增长,每个实例如此;随着先验变得扩散,无条件如此。对于任意失真,它保持与最优权重的竞争力,其比率由算法可测量的几何非相干性控制。在真实扫描数据按站点分割且各站点拥有自身校准不当的检测器时,它达到预言机超额风险;在一个将偏差与噪声分离的伴随联邦中,按噪声方差加权比完全不加权更差,而CROWD则不然。

英文摘要

A federated objective is a weighted sum of client risks, and the weights are almost always fixed in advance. We treat them instead as the only instrument of a wisdom-of-crowds mechanism: clients are noisy views of one truth, each seeing it through an independent distortion that is unbiased across the crowd. That the optimal weights are inversely proportional to the clients' error energies is classical; we begin at the question that answer presupposes, which energies belong there and whether a crowd can recover them from itself. Excess risk on the truth is of the exact order of the aggregate bias energy, so no optimizer can repair a bad weight vector; the truth itself is identifiable only up to a linear tilt, so subtracting estimated client biases provably reproduces uniform weighting. The expected per-client second moments, however, are exactly identified from the law of the crowd's disagreement by a well-conditioned linear inversion, a step random-effects meta-analysis cannot take because a source reports once; their realized counterparts are estimable up to an incoherence floor the algorithm can measure. This yields CROWD, which reads the disagreement off the optimization trajectory at no extra cost and matches a Bayesian minimax lower bound in the same constant: per instance as the horizon grows, and unconditionally as the prior becomes diffuse. For arbitrary distortions it stays competitive with the optimal weights, at a ratio governed by a geometric incoherence the algorithm can measure. On real scans split into sites with their own miscalibrated detectors it attains the oracle excess risk; on a companion federation that pulls bias and noise apart, weighting by noise variance is worse than not weighting at all, and CROWD is not.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑