arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24718cs.LG

用于优化胰腺癌治疗策略更新的联邦人工智能框架

A Federated Artificial Intelligence Framework for Optimizing Pancreatic Cancer Treatment - Strategy Update

  • University Medical Center Göttingen(哥廷根大学医学中心)
  • University of Göttingen(哥廷根大学)
  • Justus-Liebig University(吉森尤斯图斯-李比希大学)
  • Technical University of Munich(慕尼黑工业大学)
  • TUM School of Medicine and Health(慕尼黑工业大学医学与健康学院)
  • Technical University of Denmark(丹麦技术大学)
  • Max Planck Institute for Biology of Ageing(马克斯·普朗克衰老生物学研究所)
  • Philipps-University Marburg(马尔堡菲利普斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Anne-Christin Hauschild, Amirreza Aleyasin, Nils H. Beyer, Lisa Fricke, Jonas Hügel, Maryam Moradpour, Anh-Tien Nguyen, Youngjun Park, Sophia Rheinländer, Tim B… 展开作者

Anne-Christin Hauschild, Amirreza Aleyasin, Nils H. Beyer, Lisa Fricke, Jonas Hügel, Maryam Moradpour, Anh-Tien Nguyen, Youngjun Park, Sophia Rheinländer, Tim Beissbarth, Elisabeth Hessmann, Martin Middeke, Matthias Lauth, Maximilian Reichert, Ulrich Sax

AI总结:

本研究提出并更新了一个联邦学习AI框架,用于优化胰腺癌治疗,通过本地Docker容器和新型FL算法处理部分重叠特征,初步结果显示可行,但大规模部署仍具挑战。

AI中文摘要:

虽然理论上,一种涉及患者同意集中收集和分析数据的中心化方法能够提供最佳的数据质量和预测性能,但在实践中并不总是可行的。联邦学习(FL)架构已被证明是一种在GDPR边界内使用和访问分布式疾病相关资源的非常有前景的方法。在之前的一篇病例报告中,我们描述了参与站点的前提条件,以及为改善胰腺癌亚型识别和评估治疗方案而准备数据、人员和基础设施所需的管理和流程相关步骤。我们更新了这份报告,分享了我们在应对挑战方面的经验,并展示了实际联邦学习AI管道的初步结果。在参与站点,我们必须在本地FL中心(在我们的案例中,是一个集中开发并分布式部署的Docker容器)进行数据提取和转换后,识别并注释可访问的数据。该容器包含生成本地模型的FL脚本。我们应用了一种新开发的FL算法,该算法考虑了所有本地特征,包括特定于本地站点的部分重叠特征。理论上,在癌症环境中,使用德国肿瘤学核心数据集(oBDS)进行注释应该能够成功,该数据集已被用于向癌症登记处的强制性报告,并且可以在FL环境中持续使用。正如我们用公共数据集所展示的,FL算法能够稳健地处理部分重叠的特征。主要障碍,包括理顺基础设施的运营概念、为此类新型架构获得伦理批准以及对每个站点的支持,都已被解决。然而,未来扩展这种方法面临障碍;虽然纳入更广泛的多模态数据集应该是可行的,但大规模部署到更多站点仍然具有挑战性。

英文摘要:

While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predictive performance, it is not always feasible in practice. Federated Learning (FL) architectures have shown to be a very promising approach to use and access distributed disease related resources within the GDPR boundaries. In a previous case report, we described the preconditions at the participating sites and necessary administrative and process related steps to prepare data, people and infrastructure for improving subtype identification and assessing treatment options in pancreatic cancer. We update this report sharing our experience in tackling the challenges and show preliminary results of the actual federated learning AI pipelines. At the participating sites, we have to identify and annotate the data being accessible after extraction and transformation in a local FL hub - in our case a centrally developed and distributively deployed Docker container. This container comprises the FL scripts generating local models. We apply a newly developed FL algorithm considering all local features, including partial overlapping features specific to the local sites. Theoretically, an annotation in a cancer setting should succeed using the German oncology core data set (oBDS), which is already utilized for mandatory reporting to cancer registries, and can be sustained in the FL setting. The FL algorithms deal robustly with partially overlapping features as we showed with public data sets. Major roadblocks including straightening operational concepts for the infrastructures, ethics approval for such novel architectures and support for every site have been addressed. However, scaling up this approach in the future faces hurdles; while including broader multi-modal data sets should be feasible, large-scale deployment to more sites remains challenging.

补充信息

↑