arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25114cs.LGcs.AI

Flower Hub:用于联邦学习仿真与部署的可复现基准测试平台

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tava… 展开作者

Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane

首次发表
浏览论文内容

中文总结 AI 辅助

Flower Hub是一款联邦学习基准测试平台,可将基准打包为可执行应用,支持跨仿真与部署运行,涵盖多领域任务,推动联邦学习基准测试向可复用方向发展。

中文摘要 AI 辅助

联邦学习(FL)已成为在分散式数据上训练模型的关键方法,但联邦学习中的基准测试仍难以复现、比较和扩展。现有评估通常与自定义基础设施绑定,发布的研究代码不完整,且主要在仿真环境中进行,这限制了其可移植性和实际相关性。我们推出Flower Hub,这是一个用于发布、发现和执行去中心化及联邦应用的平台。我们展示了它如何通过将基准测试打包为可执行、带版本的应用,并附带标准化元数据、固定依赖项和明确的评估工作流,从而实现可复现的基准测试。我们用一个跨域基准测试套件实例化该方法,涵盖跨孤岛和跨设备场景,包括医学成像、金融表格学习、法律指令微调、钓鱼URL检测和音频标注等任务。我们进一步证明,同一基准测试应用可在仿真和部署运行时中运行,无需更改应用代码,从而能在不同学习环境中进行统一评估。除模型质量外,我们的基准设计还支持感知系统的报告,包括运行时和通信指标。这项工作将联邦学习场景中的基准测试从临时代码构件推进到可移植、可执行且可重复使用的基准测试应用。

英文摘要

Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to reproduce, compare, and extend. Existing evaluations are often tied to custom infrastructure, released as incomplete research code, and conducted primarily in simulation, which limits portability and practical relevance. We present Flower Hub, a platform for publishing, discovering, and executing decentralized and federated applications. We show how it enables reproducible benchmarking by packaging benchmarks as executable, versioned applications with standardized metadata, pinned dependencies, and explicit evaluation workflows. We instantiate this approach with a multi-domain benchmark suite spanning cross-silo and cross-device settings, and including tasks in medical imaging, financial tabular learning, legal instruction tuning, phishing URL detection, and audio tagging. We further demonstrate that the same benchmarking application can run across both simulation and deployment runtimes without changing the application code, enabling unified evaluation across varying learning environments. Beyond model quality, our benchmark design supports system-aware reporting, including runtime and communication metrics. This work advances benchmarking in FL settings from ad hoc code artifacts towards portable, executable, and reusable benchmark applications.

发表机构

  • Flower Labs
  • University of Cambridge(剑桥大学)
  • University of Auckland(奥克兰大学)
  • University of Melbourne(墨尔本大学)
  • Vector Institute
  • Fraunhofer IMS(弗劳恩霍夫应用固体物理与材料力学研究所)
  • NetCompany
  • Gachon University(嘉泉大学)
  • Owkin
  • Sony AI(索尼人工智能)

机构由 AI 辅助整理,请以论文原文为准。

↑