arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15516cs.LG

UniFed-VLM:面向多异质性视觉语言模型的联邦指令调优

UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity

Pengyu Wang, Baochen Xiong, Xiaoshan Yang, Yifan Xu, Zhang Qimeng, Haifeng Chen, Changsheng Xu

首次发表
浏览论文内容

中文总结 AI 辅助

UniFed-VLM是解决任务、模态、模型架构联合异质性的VLM联邦指令调优框架,含FedCSA与TCoD组件,在多基准数据集上优于现有FL方法。

中文摘要 AI 辅助

视觉语言模型(Vision-Language Models, VLMs)在多模态理解与生成任务中展现出强大性能,但其微调通常依赖中心化数据,这在医疗等领域引发隐私担忧。联邦学习(Federated Learning, FL)可在不共享原始数据的情况下实现模型训练,成为自然解决方案,但将FL应用于VLM指令调优极具挑战性:VLMs参数规模庞大,且真实场景中客户端在任务、模态、模型架构方面存在显著异质性,现有方法多聚焦于简化设置,无法处理此类多维异质性场景。本研究针对任务、模态、模型架构联合异质性下的联邦指令调优问题,提出UniFed-VLM——一种解决多类型异质性的VLM统一联邦指令调优框架,包含两大核心组件:1)联邦补偿子空间聚合(Federated Compensated Subspace Aggregation, FedCSA),通过动态加权与补偿对参数高效适配器进行子空间对齐聚合,以缓解异质性引发的冲突;2)两阶段协作蒸馏(Two-stage Collaborative Distillation, TCoD),通过互蒸馏适配器(Mutual Distillation Adapter, MDA)与基于混合专家的蒸馏策略,实现异构模型间的有效知识迁移。在多个基准数据集上开展实验,结果表明,与现有FL方法相比,UniFed-VLM在各类任务中实现了更强的平均性能,源代码可访问:this https URL。

英文摘要

Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine-tuning of VLMs typically relies on centralized data, which raises privacy concerns in certain domains (e.g. healthcare). Federated Learning (FL) provides a natural solution by enabling model training without sharing raw data. However, applying FL to VLM instruction tuning is highly challenging. VLMs have substantial parameter scales, and in real-world scenarios, clients exhibit significant heterogeneity in tasks, modalities, and model architectures. Existing methods mainly focus on simplified settings and are unable to handle such multi-dimensional heterogeneous scenarios. In this work, we study federated instruction tuning under joint heterogeneity in tasks, modalities, and model architectures. We propose UniFed-VLM, a unified federated instruction tuning framework for VLMs that addresses multiple types of heterogeneity. It consists of two key components: 1) Federated Compensated Subspace Aggregation (FedCSA), which performs subspace-aligned aggregation of parameter-efficient adapters with dynamic weighting and compensation to mitigate heterogeneity-induced conflicts; 2) Two-stage Collaborative Distillation (TCoD), which enables effective knowledge transfer across heterogeneous models via a Mutual Distillation Adapter (MDA) and a mixture-of-experts-based distillation strategy. We conduct experiments on multiple benchmark datasets, and the results show that UniFed-VLM achieves stronger average performance across diverse tasks compared with existing FL methods. The source code is available at: https://github.com/wangpengyu2004/UniFed-VLM.

↑