arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FM$^2$:用于异构多模态医学成像的统一联邦基础模型

FM$^2$: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

Shengchao Chen, Ting Shu

arXiv 2607.13386首次发表:更新:

发表机构

School of Artificial Intelligence, Shenzhen University; Australian AI Institute, University of Technology Sydney(深圳大学人工智能学院; 悉尼科技大学澳大利亚人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对医学成像基础模型构建中隐私与任务统一问题,提出FM$^2$框架,通过从头训练核心主干、结合预训练编码器、配备双混合专家模块及正则化器,并引入字幕增强学习,实现跨模态泛化,优于现有联邦基线。

AI 中文摘要

构建医学成像基础模型需要跨机构汇总数据,但隐私法规禁止集中聚合。现有联邦基础模型存在要么对医学领域迁移性差的自然图像模型进行微调,要么在单一模态内从头训练,缺乏统一任务灵活性的问题。本文识别出成像模态异质性挑战,客户端存在重叠和非重叠两种结构模式。提出FM$^2$统一框架,从头训练核心主干以保持医学领域保真度,可结合生物医学预训练编码器进行视觉语言对齐。配备双混合专家模块及异构模态对齐正则化器,具有收敛和泛化保证。还纳入字幕增强学习,实验证实其优于现有联邦基线且跨模态泛化能力强。

英文摘要

Building foundation models for medical imaging requires pooling data across institutions, yet privacy regulations prohibit centralized aggregation. Existing Federated Foundation Models either fine-tune natural-image models with poor medical-domain transfer, or train from scratch within a single modality, lacking the flexibility to unify tasks. We identify an under-explored challenge, Imaging Modality Heterogeneity, where clients operate under two structural regimes: Overlapped (shared modalities with heterogeneous label distributions) and Non-overlapped (fully disjoint modalities per client). We propose FM$^2$, a unified framework that trains the core backbone from scratch to preserve medical domain fidelity while optionally incorporating biomedical pretrained encoders for vision-language alignment. FM$^2$ equips each client with dual Mixture-of-Experts modules (a Class-wise MoE for personalized category knowledge and a Domain-wise MoE for shared cross-modality representations), coupled with a Heterogeneous Modality Alignment (HMA) regularizer that explicitly aligns modality-specific expert parameters, admitting provable $O(1/\sqrt{T})$ convergence and generalization guarantees. FM$^2$ further incorporates Caption-Enhanced Learning (CEL), where locally retained GPT-4o-generated captions serve as a textual semantic bridge enabling representation transfer across clients with disjoint modalities, and demonstrates extensibility to Federated Medical VQA. Experiments on our MIMH benchmark (classification and CEL) and real-world medical VQA datasets confirm consistent superiority over state-of-the-art federated baselines and strong out-of-modality generalization across all three tasks.

CommentsAccepted by ACM MM 2026 (Main Track): the 34th ACM International Conference on Multimedia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑