SynPre-FL:合成数据驱动的预训练集成联邦学习训练框架
SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework
浏览论文内容
中文总结 AI 辅助
研究针对联邦学习在临床风险预测中面临的问题及局限,提出SynPre-FL框架,结合合成EHR生成与合成预训练联邦学习,通过潜在自动编码器扩散模型等技术,提升非IID条件下预测的稳健性、可扩展性与可解释性。
中文摘要 AI 辅助
联邦学习为隐私保护临床风险预测提供了一种有前景的方法,但受限于数据共享受限、客户端异质性、类别不平衡以及缺乏真实表格电子健康记录(EHR)基准,其部署仍然有限。合成数据生成可缓解数据稀缺问题,但其与联邦优化的集成尚未得到充分系统研究。我们提出了SynPre-FL,这是一个统一框架,将高保真合成EHR生成与合成预训练联邦学习相结合,以在非IID条件下进行稳健预测。潜在自动编码器扩散模型生成隐私保护合成队列,用于预热联邦训练。随后使用类别平衡局部目标、近端正则化和自适应服务器聚合进行异质性感知优化。事后校准和联邦安全可解释性支持可靠且可解释的风险估计。实验表明,合成生成器保留了单变量、双变量和多变量结构,同时抵御成员推理和重建攻击。生成的数据在TSTR、TRTS和基于模型的评估中具有强大的下游效用。在具有5、10和15个异构客户端的联邦设置中,SynPre-FL始终比基线方法提高鲁棒性和可扩展性,特别是在严重的非IID碎片化情况下。校准提高了概率可靠性,而SHAP分析在不同联邦规模下产生稳定且临床相关的特征归因。因此,SynPre-FL为将合成数据与联邦学习相结合提供了一个实用且可重复的框架,以实现从分布式表格EHR数据进行隐私感知、可解释且稳健的临床预测。
英文摘要
Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks. Synthetic data generation may alleviate data scarcity, yet its integration with federated optimisation has received limited systematic study. We propose SynPre-FL, a unified framework combining high-fidelity synthetic EHR generation with synthetic-pretrained FL for robust prediction under non-IID conditions. A latent autoencoder-diffusion model generates privacy-preserving synthetic cohorts, which are used to warm-start federated training. This pretraining is followed by heterogeneity-aware optimisation using class-balanced local objectives, proximal regularisation, and adaptive server aggregation. Post-hoc calibration and federated-safe explainability support reliable and interpretable risk estimates. Experiments show that the synthetic generator preserves univariate, bivariate, and multivariate structure while protecting against membership-inference and reconstruction attacks. The generated data achieve strong downstream utility under TSTR, TRTS, and model-based evaluations. Across federated settings with 5, 10, and 15 heterogeneous clients, SynPre-FL consistently improves robustness and scalability over baseline methods, especially under severe non-IID fragmentation. Calibration improves probability reliability, while SHAP analysis produces stable and clinically coherent feature attributions across federation sizes. SynPre-FL therefore provides a practical and reproducible framework for combining synthetic data with FL to enable privacy-aware, interpretable, and robust clinical prediction from distributed tabular EHR data.