arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FedTVD:面向鲁棒联邦学习的数据质量与数量平衡方法

FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning

Radwan Selo, Majid Kundroo, Taehong Kim

arXiv 2608.09221首次发表:更新:

发表机构

School of Information and Communication Engineering, Chungbuk National University(忠北国立大学信息与通信工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FedTVD是一种联邦学习算法,通过结合总变差距离与数据集规模进行双重加权,缓解数据不平衡问题,在多数据集及异质性水平下性能优于FedAvg等现有方法。

AI 中文摘要

联邦学习(FL)支持在分布式客户端设备间开展协作式模型训练,同时保障数据隐私。但FL面临数据异质性带来的重大挑战,尤其是标签分布偏斜和数据集规模差异问题,这会引发模型更新偏差并阻碍收敛。为解决该问题,我们提出FedTVD,一种新型FL算法,在聚合时通过同时考量数据质量与数量来为客户端贡献赋予权重。与仅依赖数据集规模进行客户端加权的传统FL方法(如FedAvg)不同,FedTVD整合总变差距离(TVD)来衡量每个客户端的局部标签分布与均匀全局分布的差异,标签分布高度偏斜的客户端会被赋予更低权重,避免存在不平衡问题的数据集对全局模型产生过度影响;同时纳入数据集规模以确保可扩展性与公平性。这种双重加权机制有效缓解了数据不平衡的影响,可得到更稳定、泛化能力更强的全局模型。实验结果表明,FedTVD在所有数据集(FMNIST、CIFAR-10、CIFAR-100)和所有数据异质性水平下,均持续优于现有最优方法;值得注意的是,在高度偏斜数据的CIFAR-10上,它相比FedAvg实现了最高达10.6%的性能提升,且在中度和独立同分布(IID)设置下仍保持顶尖性能。

英文摘要

Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and variations in dataset sizes, which can lead to biased model updates and hinder convergence. To address this, we propose FedTVD, a novel FL algorithm that weights client contributions during aggregation by considering both data quality and quantity. Unlike traditional FL approaches such as FedAvg, which rely solely on dataset size for client weighting, FedTVD integrates Total Variation Distance (TVD) to measure the divergence between each client's local label distribution and a uniform global distribution. Clients with highly skewed distributions receive lower weights, preventing unbalanced datasets with imbalances from disproportionately influencing the global model. At the same time, dataset size is incorporated to ensure scalability and fairness. This dual-weighting mechanism effectively mitigates the impact of data imbalance, leading to more stable and generalized global models. Experimental results show that FedTVD consistently outperforms state-of-the-art methods across all datasets (FMNIST, CIFAR-10, and CIFAR-100) and all levels of data heterogeneity. Notably, it achieves up to 10.6% improvement over FedAvg on CIFAR-10 under highly skewed data, while maintaining top performance even under moderate and IID settings.

Journal refFuture Generation Computer Systems, Volume 176, March 2026, 108177

DOI:10.1016/j.future.2025.108177

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑