arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38925cs.AIcs.LG

原型引导的双边对齐多模态联邦学习

Prototype-guided Bilateral Alignment Multimodal Federated Learning

  • Hong Kong Baptist University(香港浸会大学)
  • Sun Yat-Sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianchi Liao Tianchi_Liao, Lele Fu, Sheng Huang, Qing Hu, Hong-Ning Dai, Chuan Chen

AI总结:

针对多模态联邦学习中模型异构与模态不平衡问题,提出MFedPBA框架,通过特征级投影编码器对齐和决策级熵加权logit原型聚合,显著优于现有基线。

AI中文摘要:

多模态联邦学习(MFL)已成为利用分布式数据提升模型性能的关键范式。然而,现有方法主要依赖于模型同质性和模态分布均衡的理想化假设,使其难以适应以异构客户端架构和严重模态不平衡为特征的实际场景。为应对这些挑战,我们提出了一个多模态联邦学习原型引导的双边对齐(MFedPBA)框架。MFedPBA通过双重对齐机制实现稳健的知识协同:(i)在特征层面,它通过由对比学习和Gromov-Wasserstein距离优化的投影编码器对齐异构特征空间;(ii)在决策层面,它采用熵加权聚合自然对齐的logit原型。这一新颖设计通过联合处理异构特征空间和集体聚合决策,实现了稳健的MFL。大量实验表明,在模型异构和模态不平衡条件下,我们的方法显著优于最先进的基线方法。

英文摘要:

Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing methods predominantly rely on idealized assumptions of model homogeneity and balanced modality distributions, rendering them ill-suited for practical scenarios characterized by heterogeneous client architectures and severe modality imbalance. To address these challenges, we propose a \textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment (MFedPBA) framework. MFedPBA facilitates robust knowledge synergy through a dual alignment mechanism: (i) at the feature level, it aligns heterogeneous feature spaces via a projection encoder optimized by contrastive learning and the Gromov-Wasserstein distance; (ii) at the decision level, it employs an entropy-weighted aggregation of naturally aligned logit prototypes. This novel design achieves robust MFL by jointly tackling heterogeneous feature spaces and collectively aggregating decisions. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines under conditions of model heterogeneity and modality imbalance.

补充信息

↑