发表机构
Turku University of Applied Sciences(图尔库应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出DP-SimAgg框架,结合相似度加权聚合与服务器端差分隐私,在FeTS 2022数据集上验证其可在保护隐私的同时保持脑病灶分割的竞争力性能。
AI 中文摘要
联邦学习(FL)可实现多机构在不共享敏感数据的情况下协同训练机器学习模型,尤其适用于医学影像应用,但各机构间异构数据分布以及模型更新可能存在的信息泄露仍是重要挑战。本研究提出DP-SimAgg,一种结合相似度加权聚合与服务器端差分隐私机制的隐私保护联邦学习框架:该方法对参与者更新应用L2裁剪以限制范围,计算基于相似度的聚合权重以缓解非独立同分布(non-IID)数据分布的影响,并在中央服务器注入校准高斯噪声,在假设的敏感度边界下提供每轮隐私保障。该框架基于英特尔的OpenFL平台实现,在FeTS 2022数据集上评估,该数据集包含1251例用于脑肿瘤分割的多模态MRI扫描。实验结果表明,DP-SimAgg在提供隐私保护的同时保持了有竞争力的分割性能:在严格的每轮隐私预算(ε=1,20轮累计ε_total=20)下,该方法对增强肿瘤(ET)、肿瘤核心(TC)和全肿瘤(WT)区域的Dice分数分别达到0.6357、0.5305和0.5274;在更宽松的每轮预算(ε=10,累计ε_total=200)下,性能接近非隐私基线,同时结合中央高斯机制,在假设的敏感度边界下实现每轮(ε,δ)-DP保障。这些结果凸显了DP-SimAgg在医学影像应用中实现隐私保护协同学习的潜力。
英文摘要
Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making it particularly suitable for medical imaging applications. However, heterogeneous data distributions across institutions and potential information leakage through model updates remain important challenges. In this work, we propose DP-SimAgg, a privacy-preserving federated learning framework that integrates similarity-weighted aggregation with a server-side differential privacy mechanism. The proposed method applies L2 clipping to bound collaborator updates, computes similarity-based aggregation weights to mitigate the effects of non-IID data distributions, and injects calibrated Gaussian noise at the central server, providing per-round privacy guarantees under the assumed sensitivity bound. The framework is implemented using Intel's OpenFL platform and evaluated on the FeTS 2022 dataset consisting of 1251 multi-modal MRI scans for brain tumor segmentation. Experimental results demonstrate that DP-SimAgg maintains competitive segmentation performance while providing privacy protection. Under a strict per-round privacy budget (epsilon = 1, cumulative epsilon_total = 20 over 20 rounds), the method achieves Dice scores of 0.6357, 0.5305, and 0.5274 for the enhancing tumor (ET), tumor core (TC), and whole tumor (WT) regions, respectively. With a more relaxed per-round budget (epsilon = 10, cumulative epsilon_total = 200), performance approaches that of the non-private baseline while incorporating a central Gaussian mechanism with per-round (epsilon, delta)-DP accounting under the assumed sensitivity bound. These results highlight the potential of DP-SimAgg for enabling privacy-preserving collaborative learning in medical imaging applications.