arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多光谱与SAR图像理解的通用视觉语言模型(VLM)的轻量适配

Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding

Shanji Liu, Kelu Yao, Junxiao Xue, Chenghui Lv, Xiangyang Miao, Yekai Huang, Yaying Chen, Chao Li

arXiv 2609.02187首次发表:更新:

发表机构

Zhejiang University; Zhejiang Lab(浙江大学; 之江实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出通过渲染多光谱与SAR输入视图并采用LoRA轻量适配,将通用VLMs迁移至多光谱和SAR图像理解任务,在土地覆盖识别等任务上取得良好效果,且可跨架构与数据集迁移。

AI 中文摘要

通用视觉语言模型(VLMs)如今具备强大的视觉识别、指令跟随与生成能力。然而,多数预训练视觉编码器围绕三通道自然图像构建,无法直接适配原生多光谱测量或合成孔径雷达(SAR)等观测数据。将VLMs适配至这些传感器通常需要专用编码器与领域预训练,这会降低复用更强通用预训练检查点的效率。本文表明,通用VLMs的多图像接口提供了一种轻量替代方案:我们的协议将每个观测数据渲染为5个光学视图与1个SAR视图,在提示词中命名这些视图,并通过LoRA(低秩适配)方法适配语言网络与选定的视觉Transformer块。这一方式通过现有视觉接口呈现了波段合成、光谱指数与雷达后向散射信息。针对土地覆盖识别任务,结构化监督将预测类别与传感器证据关联;我们还构建了偏好对,其中真实标签被省略但保留其支撑证据,以鼓励生成与观测一致的完整预测。在BigEarthNet-v2衍生的平衡六类土地覆盖基准上,适配后的Qwen3-VL达到了0.8275的微F1值。相同输入与适配协议提升了所有4种测试的VLM架构的性能,并可迁移至Sen1Floods11洪水验证及this http URL的图像描述任务。图像移除与不匹配控制实验表明,适配后的模型确实使用了所提供的传感器观测数据。综上,这些结果证明,VLMs可通过渲染输入与紧凑的LoRA适配,被重新用于多光谱与SAR任务,而无需训练新的基础模型。

英文摘要

General-purpose vision-language models (VLMs) now support strong visual recognition, instruction following, and generation. However, most pretrained visual encoders are built around three-channel natural images and do not directly accommodate observations such as native multispectral measurements or synthetic aperture radar (SAR). Adapting VLMs to these sensors typically requires dedicated encoders and domain pretraining, slowing the reuse of stronger general-purpose checkpoints. We show that the multi-image interface of general-purpose VLMs offers a lightweight alternative. Our protocol renders each observation as five optical views and one SAR view, names them in the prompt, and adapts the language network and selected visual transformer blocks with LoRA. This exposes band composites, spectral indices, and radar backscatter through an existing visual interface. For land-cover recognition, structured supervision couples predicted classes with sensor evidence. We further construct preference pairs in which a true label is omitted while its supporting evidence is retained, encouraging complete predictions that remain consistent with the observations. On a balanced six-class land-cover benchmark derived from BigEarthNet-v2, the adapted Qwen3-VL reaches 0.8275 micro F1. The same input and adaptation protocol improves all four tested VLM architectures and transfers to Sen1Floods11 flood verification and BigEarthNet.txt captioning. Image removal and mismatch controls show that the adapted models use the supplied sensor observations. Together, these results demonstrate that VLMs can be repurposed for multispectral and SAR tasks through rendered inputs and compact LoRA adaptation, without training a new foundation model.

Comments18 pages, 7 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑