发表机构
The University of British Columbia; Fudan University; Royal College of Science, Imperial College London; Dyson School of Design Engineering; Suzhou Institute of Biomedical Engineering and Technology (SIBET), Chinese Academy of Sciences; East China Normal University(英属哥伦比亚大学; 复旦大学; 伦敦帝国理工学院皇家科学学院; 戴森设计工程学院; 中国科学院苏州生物医学工程技术研究所; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对联邦多模态大语言模型微调中灾难性遗忘问题,提出FedCMM框架,在参数、数据、聚合三个层面嵌入持续学习保障,经实验验证该框架在准确性和反向迁移上优于基线,能实现跨异构网络AI部署的稳健进化适应。
AI 中文摘要
跨分布式网络对多模态大语言模型(MLLM)进行联邦微调,可在保护隐私的情况下适应不断变化的数据流,但灾难性遗忘阻碍其在动态环境中的稳健部署。为应对这一挑战,我们提出了联邦持续多模态学习(FedCMM)框架,它在三个互补层面将持续学习保障嵌入联邦优化循环。在参数层面,模态感知弹性权重整合为视觉编码器、语言主干和跨模态投影仪计算单独的Fisher信息矩阵;在数据层面,每个客户端训练一个轻量级本地生成重放模块来合成无原始数据的嵌入级多模态重放元组;在聚合层面,任务相似性感知梯度聚合通过梯度余弦相似性自主过滤和重新加权客户端更新。实验表明,FedCMM在准确性和反向迁移方面优于近期基线,证实了整体、模态感知优化可实现跨异构网络AI部署的稳健进化适应。
英文摘要
Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations. Addressing this challenge is especially critical for autonomous networked AI operating in safety-sensitive domains, such as content moderation, where reliable retention of prior knowledge underpins system integrity. To overcome this, we propose Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels. At the parameter level, modality-aware elastic weight consolidation computes separate Fisher information matrices for the vision encoder, language backbone, and cross-modal projector, providing granular, asymmetry-aware protection against modality-specific forgetting. At the data level, each client trains a lightweight local generative replay module to synthesize raw-data-free embedding-level multimodal replay tuples without any raw data sharing. At the aggregation level, Task-similarity-aware gradient aggregation autonomously filters and reweights client updates by gradient cosine similarity, suppressing conflicting directions and stabilizing the global learning trajectory. Extensive experiments on two benchmarks demonstrate that FedCMM consistently outperforms recent baselines on accuracy and backward transfer, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.
CommentsAccepted by IEEE JSTSP
DOI:10.1109/JSTSP.2026.3740619