AI 中文总结
针对联邦学习中客户端资源异构导致的不平衡问题,提出FEAST框架,通过子超网络路由和稀疏聚合联合训练多子网络,在多数据集上实现更高准确率并降低参数流量。
AI 中文摘要
联邦学习(FL)需要服务于计算能力各异的设备。固定模型无法适配所有设备,而每次部署训练一个模型的成本过高。联邦超网络训练则学习一个弹性模型,该模型包含不同规模的子网络,随后为每个设备部署合适的子网络。然而,当客户端推理预算存在差异时,高成本子网络独有的参数可被更少的客户端访问。我们提出FEAST,这是一种联邦共享空间训练框架,通过在每个客户端的预算限制内联合训练多个子网络来缓解这种不平衡。定制预算的子超网络路由仅发送超网络的相关部分,而稀疏聚合则合并返回的参数切片。训练后的超网络可直接服务于联邦期间使用的子网络,并且支持在无需联邦重训练的情况下事后提取额外的子网络。我们进一步表明,在异构联邦学习模拟中,独立分配客户端的训练数据量和推理预算会扭曲准确率-推理成本的比较,为此我们引入了单参数γ分配协议来控制这种耦合关系。在我们的实验设置中,SuperFedNAS和DeepFedNAS超网络训练过程在25M推理MACs时接近随机水平,在596M推理MACs时最多达到17.09%;FEAST在596M推理MACs时达到71.06%,比最强的模型异构权重共享基线在其最大层级的准确率高出2.4个百分点。在CIFAR-100、CINIC-10和TinyImageNet-200数据集上,当每个客户端接收其最大可负担的子网络时,FEAST在所有评估的权重共享方法中实现了最高的人口平均准确率。与全超网络传输相比,子超网络路由将聚合模型参数流量降低了6.8倍。
英文摘要
Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per deployment limit is costly. Federated supernet training instead learns one elastic model with differently sized subnetworks, then deploys a suitable one to each device. When client inference budgets differ, however, parameters exclusive to high-cost subnetworks are reachable by fewer clients. We propose FEAST, a federated shared-space training framework that counters this imbalance by jointly training multiple subnetworks within each client's limit. Budget-tailored sub-supernet routing sends only the relevant supernet portion, and sparse aggregation merges the returned parameter slices. The trained supernet directly serves the subnetworks used during federation and supports post-hoc extraction of additional subnetworks without federated retraining. We further show that independently assigning clients' training-data volumes and inference budgets can distort accuracy--inference-cost comparisons in heterogeneous FL simulations, and introduce a one-parameter $γ$-allocation protocol to control this coupling. In our experimental setup, the SuperFedNAS and DeepFedNAS supernet training procedures remain near chance at 25M and reach at most $17.09\%$ at $596$M inference MACs; FEAST reaches $71.06\%$ at $596$M, $2.4$ points above the strongest model-heterogeneous weight-sharing baseline at its largest tier. Across CIFAR-100, CINIC-10, and TinyImageNet-200, FEAST achieves the highest population-averaged accuracy among the evaluated weight-sharing methods when each client receives its largest affordable subnetwork. Sub-supernet routing reduces aggregate model-parameter traffic by $6.8\times$ relative to full-supernet transmission.