评估不完美数据集对联邦学习中客户端选择的影响
Assessing the Impacts of Imperfect Datasets on Client Selections in Federated Learning
浏览论文内容
中文总结 AI 辅助
该研究针对联邦学习中不完美数据集引发的问题,提出隐私保护的客户端贡献评分方法,实验验证了评估效果。
中文摘要 AI 辅助
联邦学习(FL)是一种流行的分布式学习框架,多个客户端执行本地训练,服务器聚合本地更新的模型,FL在实现去中心化训练的同时保护客户端数据集的隐私。然而,非独立同分布(non-IID)或有噪声的数据集可能导致模型准确率低或收敛延迟高,通过客户端选择排除这些客户端可缓解问题,但严重偏向的客户端选择也会降低学习性能。本研究首先通过实验测量非IID数据(包括数据量和标签分布的偏斜)、有噪声数据以及客户端选择的公平性对模型准确率和收敛性的影响;随后提出一种隐私保护的评分方法来评估FL中每个客户端的贡献,并通过实验证明所提评估方法的有效性。
英文摘要
Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models. FL enables decentralized training while preserving the privacy of clients' datasets. However, non-independent and identically distributed (non-IID) or noisy datasets can lead to low model accuracy or high convergence latency. Precluding these clients through client selection may mitigate the problem, but heavily biased client selections may also degrade the learning performance. In this study, we first experimentally measure the impact of non-IID data (including skews in data quantity and label distribution), noisy data, and fairness in client selection on model accuracy and convergence. We then propose a privacy-preserving scoring method to assess each client's contribution in FL, with experiments conducted to demonstrate the effectiveness of the proposed assessment.
发表机构
- National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。