arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08219cs.CVcs.LG

多器官图像上联邦学习的基准评估

Benchmark Evaluation of Federated Learning on Multi-organ Images

  • Hunan Provincial Key Laboratory on Bioinformatics, School of Computer Science and Engineering, Central South University(湖南生物信息省重点实验室,计算机科学与工程学院,中南大学)
  • Xinjiang Engineering Research Center of Big Data and Intelligent Software, School of software, Xinjiang University(新疆大数据与智能软件工程研究中心,软件学院,新疆大学)

机构由 AI 辅助整理,请以论文原文为准。

Junbin Mao, Xu Tian, Jianchun Zhu, Ludi Li, Jin Liu

AI总结:

针对医学数据隐私及异质性致联邦学习性能评估难的问题,开发MobenFL基准,集成20种前沿算法和22个数据集,涵盖12个器官,评估维度更丰富,能全面深入评估联邦学习在医学领域临床应用的效果。

AI中文摘要:

医学数据的隐私要求及其在器官和模态上的巨大差异阻碍了医学人工智能的临床应用。联邦学习(FL)是克服这些挑战的可行方法。由于FL算法不断涌现且医学数据高度异质性,客观评估其在实际临床环境中的性能仍很困难。因此,一个全面的联邦医学成像基准作为统一评估标准对推动该技术走向可靠临床应用至关重要。现有联邦医学成像基准未充分纳入最新算法,限于单器官或模态数据,且过度强调模型准确性。为应对这些挑战,我们开发了MobenFL基准。它集成20种前沿FL算法和22个医学成像数据集,涵盖人体12个关键器官,在广度上超越现有基准。在评估维度上,MobenFL不仅评估性能,还系统纳入算法效率和隐私保护能力等关键指标。此外,它针对涉及不同疾病、设备和成像模态的复杂实际临床场景进行专门评估,为FL在医学领域的临床应用提供了全面深入的评估框架。

英文摘要:

The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI. Federated learning (FL) is a feasible approach to overcome these challenges. Due to the continuous emergence of FL algorithms and the highly heterogeneous nature of medical data, objectively evaluating their performance in real-world clinical settings remains difficult. Therefore, a comprehensive federated medical imaging benchmark, serving as a unified evaluation standard, is crucial for advancing the technology toward reliable clinical application. Existing federated medical imaging benchmarks have not yet adequately incorporated state-of-the-art algorithms, are limited to data from single organs or modalities, and overly emphasize model accuracy, making it difficult to comprehensively assess the overall efficacy of FL in real-world medical environments. To address these challenges, we developed the MobenFL benchmark. This benchmark integrates 20 cutting-edge FL algorithms and 22 medical imaging datasets, covering 12 critical organs across the human body, surpassing existing benchmark in breadth. In terms of evaluation dimensions, MobenFL not only assesses performance but also systematically incorporates key metrics such as algorithmic efficiency and privacy protection capabilities. Additionally, it conducts specialized evaluations for complex real-world clinical scenarios involving different diseases, devices, and imaging modalities, thereby providing a comprehensive and in-depth evaluation framework for the clinical application of FL in the medical field.

↑