arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39250cs.LGcs.DC

面向计算高效同步联邦学习的客户端与训练数据选择

Client and Training Data Selection for Computationally Efficient Synchronized Federated Learning

  • Technical University of Berlin(柏林工业大学)
  • The University of Osaka(大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

Muzaffer Citir, Hiroki Nishikawa, Sangyoung Park

AI总结:

针对联邦学习中掉队客户端和非独立同分布数据导致的收敛慢问题,提出联合客户端与训练数据选择算法,在CIFAR-100上显著提升收敛速度和模型准确率。

AI中文摘要:

联邦学习(FL)是一种有前景的机器学习范式,它通过允许在不与云服务器共享原始数据的情况下进行学习来保护用户隐私。掉队客户端一直是FL的一个问题,因为它们会引入本地模型聚合的延迟,从而影响全局模型的收敛。因此,拥有一种既能确保全局模型快速收敛又能保证良好FL参与率的机制非常重要。FL中模型收敛的另一个问题是客户端之间数据非独立同分布(non-iid)。基于概率的客户端选择方法在非独立同分布数据下效果不佳,尤其是在客户端数量较少时。我们展示了此类方法失败的场景,并提出了一种联合客户端-训练数据选择算法,以促进FL模型的快速收敛。我们在CIFAR-100数据集上的实验表明,与先前能够考虑非独立同分布数据和异构计算的方法相比,FL模型的收敛性可以显著提高,并且模型准确率更高。

英文摘要:

Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server. Straggling clients have been a problem for FL as they introduce delays in aggregating the local models and hence, the convergence of the global model. Therefore, it is important to have a mechanism that ensures fast convergence of the global model as well as good FL participation rate. Another issue for the convergence of a model in FL is the non-independent and identically distributed (non-iid) data across the clients. Prior approaches based on probabilistic client selection do not work well under non-iid data especially when the number of clients is small. We show scenarios where such approaches fail and propose a joint client-training data selection algorithm for fast convergence of FL models. Our experiments on CIFAR-100 dataset show that convergence of the FL model can be significantly improved over prior works that can consider non-iid data and heterogeneous computation and higher model accuracy.

补充信息

↑