FedLTLib:联邦长尾学习的综合基准
FedLTLib: A Comprehensive Benchmark for Federated Long-Tail Learning
- Guangzhou University(广州大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对联邦学习中长尾分布导致尾部类别性能下降的问题,提出综合基准FedLTLib,整合多样数据集和13种算法,实现公平可复现的评估,推动移动计算稳健智能部署。
AI中文摘要:
受隐私保护计算需求不断增长的驱动,联邦学习(FL)取得了显著进展,成为连接移动边缘网络中分布式数据孤岛的关键技术。然而,在现实世界的移动计算环境中,数据由具有不同用户行为的异构移动设备生成,导致显著的长尾分布。与理想化的平衡数据集不同,现实世界中的数据表现出严重的不平衡,少数头部类别主导样本空间,而大量尾部类别(通常代表罕见但关键的边缘事件)极为稀缺。这种数据异质性(我们将其正式表征为“双重异质性”,指全局类别不平衡与局部统计偏斜的叠加)导致尾部类别性能严重下降,从而催生了联邦长尾学习(FL-LT)这一重要研究方向。为标准化评估并加速该领域研究,我们引入了FedLTLib,一个专为FL-LT设计的综合基准。针对先前工作中实验配置不一致和比较不公平的关键问题,FedLTLib建立了标准化评估框架。该平台不仅整合了反映移动数据特征的多样化基准数据集,还实现了13种最先进的FL算法(4种传统FL算法和9种FL-LT算法)。通过利用FedLTLib,研究人员可以在统一实验协议下对算法的鲁棒性和泛化能力进行公平且可复现的评估,最终推动移动计算生态系统中稳健智能的部署。
英文摘要:
Driven by the escalating demand for privacy-preserving computing, Federated Learning (FL) has witnessed remarkable progress, becoming a cornerstone technology for bridging distributed data silos in mobile edge networks. However, in real-world mobile computing environments, data is generated by heterogeneous mobile devices with varying user behaviors, leading to a significant Long-Tail Distribution. Unlike idealized balanced datasets, data in the wild manifests an acute imbalance where a minority of head classes dominate the sample space while a vast number of tail classes, often representing rare but critical edge-case events, are extremely scarce. This data heterogeneity, which we formally characterize as "Double Heterogeneity", referring to the superposition of global class imbalance and local statistical skew, precipitates severe performance deterioration on tail classes, thereby spurring the vital research direction of Federated Long-Tail Learning (FL-LT). To standardize evaluation and accelerate research in this field, we introduce FedLTLib, a comprehensive benchmark tailored for FL-LT. Addressing the critical issues of inconsistent experimental configurations and unfair comparisons in prior work, FedLTLib establishes a standardized evaluation framework. The platform not only incorporates diverse benchmark datasets reflecting mobile data characteristics but also implements 13 state-of-the-art FL algorithms (4 traditional FL algorithms and 9 FL-LT algorithms). By leveraging FedLTLib, researchers can perform fair and reproducible evaluations of algorithm robustness and generalization capabilities under a unified experimental protocol, ultimately advancing the deployment of robust intelligence in mobile computing ecosystems.