arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

遥感领域基于视觉语言模型的联邦学习中适配策略的有效性研究

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

Simon Lösche, Barış Büyüktaş, Mathis Adler, Angelos Zavras, Ioannis Papoutsis, Begüm Demir

arXiv 2608.04791首次发表:更新:

发表机构

BIFOLD - Berlin Institute for the Foundations of Learning and Data; Technische Universität Berlin; National Technical University of Athens; National Observatory of Athens; Harokopio University of Athens(BIFOLD - 柏林学习与数据基础研究所; 柏林工业大学; 雅典国家技术大学; 雅典国家天文台; 哈罗科皮奥雅典大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对遥感图像分类场景,首次对比全微调、LoRA等四种VLM适配策略在联邦学习中的表现,给出不同约束下的策略选择指南。

AI 中文摘要

联邦学习(FL)可在分散的图像档案中协作训练深度学习模型,无需数据集中,该范式在遥感(RS)领域尤为适用,因为法律规定、隐私顾虑和带宽限制会阻碍数据共享。然而,客户端间训练数据的异质性(即非独立同分布数据,non-IID数据)会阻碍收敛并限制聚合后全局模型的泛化能力。为缓解训练数据异质性的负面影响,视觉语言模型(VLM)可被用于联邦学习,因其可迁移的表示在分布偏移下表现出鲁棒性,但VLM的大量参数会大幅增加联邦场景下的通信开销和本地计算复杂度。因此,选择合适的VLM适配策略以平衡泛化能力与通信、计算约束至关重要。针对该问题,本文首次开展遥感图像分类场景下VLM适配策略用于联邦学习的对比研究,探究全微调、编码器特定微调、提示学习和低秩适配(LoRA)微调四种策略,并从三个维度分析:1)非IID数据下的泛化能力;2)通信开销;3)本地计算复杂度。在BigEarthNet-S2、EuroSAT、RESISC45和ImageNet上的实验表明,任务专业化、跨域泛化与效率间存在不同权衡。基于研究发现,本文给出不同操作约束下遥感图像分类联邦学习中VLM适配策略选择的指南,本研究代码公开于该https URL。

英文摘要

Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization. This paradigm is particularly relevant in remote sensing (RS), where legal regulations, privacy concerns, and bandwidth constraints restrict data sharing. However, the presence of training data heterogeneity across clients (known as non-IID data) can impede convergence and limit the generalization capability of the aggregated global model. To mitigate the adverse effects of training data heterogeneity, vision-language models (VLMs) can be leveraged in FL due to their transferable representations, which have demonstrated robustness under distribution shifts. However, their large parameter size may substantially increase communication overhead and local computational complexity in federated settings. Therefore, it is crucial to select an appropriate VLM adaptation strategy that balances the generalization ability with the communication and computational constraints. To address this issue, in this paper, we present the first comparative study of VLM adaptation strategies for FL in the context of RS image classification. We investigate full fine-tuning, encoder-specific fine-tuning, prompt learning, and low-rank adaptation (LoRA) tuning, and analyze them with respect to three criteria: 1) generalization capability under non-IID data, 2) communication overhead, and 3) local computational complexity. Experiments on BigEarthNet-S2, EuroSAT, RESISC45, and ImageNet reveal distinct trade-offs between task specialization, cross-domain generalization, and efficiency. Based on our findings, we derive a guideline for the selection of an appropriate VLM adaptation strategy in FL for RS image classification under different operational constraints. The code of this work is publicly available at https://git.tu-berlin.de/rsim/FL-RS-VLM.

CommentsAccepted at the SPIE Artificial Intelligence and Image and Signal Processing for Remote Sensing, Edinburgh, Scotland, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑