arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniMed-FL:用于临床诊断的鲁棒多模态联邦学习框架

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

Ayush Debnath, Ruelia Saha, Sudip Misra

arXiv 2609.10364首次发表:更新:

发表机构

Indian Institute of Technology Kharagpur; KTH Royal Institute of Technology; SRM University-AP(印度理工学院卡拉格普尔分校; 瑞典皇家理工学院; SRM大学安得拉邦分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出OmniMed-FL多模态联邦学习框架,融合影像与文本进行五类临床诊断,在非IID数据下评估多种策略,多模态融合优于单模态。

AI 中文摘要

临床诊断通常需要同时评估医学影像和患者记录。然而,标准机器学习算法无法共同分析这些数据类型。同时,遵守HIPAA和GDPR可能会限制敏感患者数据的集中聚合。这就在跨远程网络的视觉和文本上下文的安全融合方面留下了一个关键空白。因此,我们提出了OmniMed-FL,一项针对五类临床状况分类(正常、肺炎、COVID-19、胸腔积液、心脏肥大)的多模态联邦学习的受控系统研究。我们的代理语料库将3,000张公共胸部X光片与3,000条按类别条件生成的合成笔记配对,按类别而非按患者匹配。该框架在3到20个医院客户端的非IID Dirichlet分区下,对八种融合策略、三种初始化、四种缺失文本插补规则以及匹配的联邦基线进行了基准测试。由于所有笔记都是合成的,且配对并非患者级别,这些是描述性代理比较,而非诊断性能或部署就绪性的估计。在这些限制内,当客户端数(K=5)且严重偏斜(α=0.1)时,仅本地训练达到宏F1分数0.297,FedAvg达到0.662±0.074,FedProx达到0.737±0.085,匹配的FedMME风格一次性集成达到0.647±0.080,而我们提出的SCAFFOLD-AdamW自适应达到0.070±0.015,FedProx与FedAvg之间的0.075差距落在两次种子标准差中较宽的一个之内。在4×3网格上,标签偏斜最多损失0.27的F1,而客户数量增加近七倍最多损失0.10,同时双向数据量线性增长至K=20时的183.5 GiB。多模态融合在两个语料库上均领先,在合成语料库上得分为0.956,而文本为0.934,图像为0.664;在X光片语料库上得分为0.906,而文本为0.880,图像为0.737,其模型状态大小为仅文本的2.3倍。

英文摘要

Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis. However, standard machine learning algorithms cannot analyze these data types together. Meanwhile, compliance with HIPAA and GDPR can constrain centralized aggregation of sensitive patient data. This leaves a crucial void of secure fusion of visual and textual context across distant networks. Thus, we present OmniMed-FL, a controlled systems study of multimodal federated learning for five-class clinical condition classification (Normal, Pneumonia, COVID-19, Pleural Effusion, Cardiomegaly). Our proxy corpus pairs 3,000 public chest radiographs with 3,000 class-conditioned synthetic notes, matched by class, not by patient. The framework benchmarks eight fusion strategies, three initializations, four missing-text imputation rules, and matched federated baselines under non-IID Dirichlet partitioning across 3 to 20 hospital clients. As all notes are synthetic and pairing is not patient-level, these are descriptive proxy comparisons, not estimates of diagnostic performance or deployment readiness. Within those limits with clients ($K=5$) and severe skew ($α=0.1$), local-only training achieves a macro-F1 score of 0.297, FedAvg achieves $0.662\pm0.074$, FedProx $0.737\pm0.085$, a matched FedMME-style one-shot ensemble $0.647\pm0.080$, and our SCAFFOLD-AdamW adaptation $0.070\pm0.015$, the 0.075 FedProx-FedAvg gap falling inside the wider of the two two-seed standard deviations. Over a $4\times3$ grid, label skew costs up to 0.27 F1 whereas a near-sevenfold client increase costs at most 0.10, while bidirectional volume grows linearly to 183.5 GiB at $K=20$. Multimodal fusion leads on both corpora, scoring 0.956 against 0.934 for text and 0.664 for images on the synthetic corpus and 0.906 against 0.880 and 0.737 on the radiograph corpus, for $2.3\times$ the model state of text alone.

CommentsAccepted in IEEE Globecom 2026, E-Health

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑