发表机构
University of Isfahan(伊斯法罕大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过建立梯度反演与纠删码理论的联系,提出LT码启发的剥离攻击,在单轮FedSGD中精确恢复整个批次及标签,大幅超越先前攻击,揭示联邦学习隐私泄露被严重低估。
AI 中文摘要
联邦学习共享模型更新而非原始数据,然而这些更新可被反演以重建客户端的训练数据。解析重建攻击以闭式形式反演梯度,但随着批次增大其性能下降:先前的单轮攻击即使在攻击者完全控制网络参数的情况下,也仅能恢复大小为$100$的批次中的约一半数据,且已知上界限制了任何此类方法所能恢复的内容。我们在梯度反演与纠删码理论之间建立了联系,并利用这一联系构造了超越这些上界的攻击。我们的攻击能够从单轮FedSGD中精确恢复整个批次及其每个样本的标签,并在无需真实数据的情况下认证每次恢复。在八个图像和表格基准上,它们大幅优于先前的单轮攻击。即使是仅观察诚实训练网络的被动攻击者,也能在批次大小高达$128$时恢复ImageNet批次的$94$--$100\%$,这超过了先前单轮攻击在主动操纵模型时所能达到的效果;而在主动设置下,当批次大小为数百时,可恢复超过$90\%$的数据。这些结果表明,联邦学习的隐私泄露被低估了。
英文摘要
Federated learning shares model updates rather than raw data, yet these updates can be inverted to reconstruct the clients' training data. Analytic reconstruction attacks, which invert a gradient in closed form, degrade as the batch grows: prior single-round attacks recover only about half of a batch of size $100$ even when the attacker fully controls the network parameters, and known upper bounds limit what any such method can recover. We establish a connection between gradient inversion and the theory of erasure-correcting codes, and use it to construct attacks that exceed these bounds. Our attacks recover batches exactly, together with every sample's label, from a single FedSGD round, and certify each recovery without ground-truth data. On eight image and tabular benchmarks they outperform prior single-round attacks by a wide margin. Even a passive attacker who only observes an honestly trained network recovers $94$--$100\%$ of ImageNet batches at sizes up to $128$, more than prior single-round attacks achieve even with active manipulation of the model, and in the active setting more than $90\%$ is recovered at batch sizes of several hundred. These results show that the privacy leakage of federated learning has been underestimated.