发表机构
BRAC University; York University; University of Maryland(BRAC大学; 约克大学; 马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学数据合成中隐私保护及合成数据预测结构问题,提出FedDP-PALD框架,通过带掩码的门控多头注意力联合处理多模态数据,引入DP-PMA维持差分隐私,实验表明该框架能生成隐私合成表示,保持决策性能并抵抗成员推理。
AI 中文摘要
医学图像和生理信号为准确诊断提供了有价值的信息。开发诊断模型通常需要来自多个机构的患者数据,但严格的隐私法规限制了敏感临床记录的共享。联邦学习使多家医院无需交换原始数据就能训练共享模型。然而,现有方法存在两个问题:训练期间交换的信息会揭示是否使用了患者数据,且用于替代真实记录的合成数据往往无法保留其预测结构,这限制了临床应用。为解决此问题,我们提出了FedDP-PALD,这是一种在形式化隐私保证下用于多模态医学数据合成的隐私保护联邦潜在扩散框架。它通过带模态可用性掩码的门控多头注意力联合处理胸部X光图像和心电图(ECG)信号,即使在缺少一种模态时也依然有效。我们还引入了差分隐私原型混合聚合(DP-PMA),它在服务器上组合之前裁剪类级潜在原型并添加校准高斯噪声以维持(ε,δ)差分隐私。我们在肺炎MNIST、胸部MNIST和MIT-BIH数据集上评估了FedDP-PALD,对于隐私预算从ε = 1到ε = 8,差分隐私将汇总级攻击的AUROC从0.6229±0.0026降至0.5016至0.5093之间。在测试数据上,合成潜在训练实现了0.8993±0.0006的F1分数以及0.9057±0.0503的AUROC,接近真实潜在训练的0.9747±0.0132。这些结果表明,FedDP-PALD生成了隐私保护的合成表示,在保持有用决策性能的同时能强烈抵抗成员推理。
英文摘要
Medical images and physiological signals provide valuable information for accurate diagnosis. Developing diagnostic models often requires patient data from multiple institutions, although strict privacy regulations limit the sharing of sensitive clinical records. Federated learning enables multiple hospitals to train a shared model without exchanging raw data. However, existing methods face two problems: the information exchanged during training can reveal whether a patient's data were used, and synthetic data meant to replace real records often fail to preserve their predictive structure, which limits clinical use. To address this issue, we propose FedDP-PALD, a privacy-preserving federated latent diffusion framework for multimodal medical data synthesis under formal privacy guarantees. It jointly processes chest X-ray images and electrocardiogram (ECG) signals through gated multi-head attention with modality-availability masks, remaining effective even when a modality is missing. We also introduce Differentially Private Prototype Mixture Aggregation (DP-PMA), which clips class-level latent prototypes and adds calibrated Gaussian noise before combining them on the server to maintain $(ε, δ)$ differential privacy. We evaluate FedDP-PALD on PneumoniaMNIST, ChestMNIST, and MIT-BIH datasets, where differential privacy reduced summary-level attack AUROC from 0.6229 $\pm$ 0.0026 to between 0.5016 and 0.5093 for privacy budgets from $ε= 1$ to $ε= 8$. On the test data, synthetic-latent training achieved an F1 score of 0.8993 $\pm$ 0.0006 and an AUROC of 0.9057 $\pm$ 0.0503, close to the 0.9747 $\pm$ 0.0132 real-latent training. These results show that FedDP-PALD generates private synthetic representations that preserve useful decision performance while strongly resisting membership inference.