发表机构
School of Computer Science and Technology, Zhejiang University; Space-based Computing System Research Center, Zhejiang Lab(浙江大学计算机科学与技术学院; 之江实验室天基计算系统研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出OmniRSCLIP框架,通过SSBD方法将CLIP扩展至多源遥感数据,构建OmniRS5M数据集,在多项任务上验证了其性能。
AI 中文摘要
对比语言-图像学习(CLIP)已成为遥感视觉-语言理解的关键范式。然而,现有遥感对比学习方法大多构建于面向RGB的CLIP架构之上,难以合成利用合成孔径雷达(SAR)、多光谱成像(MSI)和高光谱成像(HSI)等异质传感器。为解决这一局限,我们提出OmniRSCLIP,一种支持多源传感器输入的端到端遥感视觉-语言建模对比学习框架。核心思路是在不破坏预训练视觉知识的前提下,将CLIP扩展至其固定RGB输入接口之外。为此,OmniRSCLIP引入光谱-空间基分解(SSBD),将任意通道适配建模为基重组问题:预训练CLIP的patch嵌入提供可迁移的空间基,而波长相关系数在受限视觉先验空间内生成传感器特定的嵌入核。该设计避免将异质传感器强制纳入固定通道输入空间,同时在统一图像-文本语义空间中对齐它们。我们进一步引入光谱-上下文感知的掩码对比学习方案,以抑制模态特定的冗余特征并增强细粒度图像-文本对齐。最后,为支持多模态训练,我们构建OmniRS5M,首个涵盖RGB、SAR、MSI和HSI的大规模遥感图像-文本语料库。在检索、零样本分类和语义定位任务上的实验表明,OmniRSCLIP在保留强大RGB域性能的同时,有效将CLIP扩展至多源异质遥感模态。
英文摘要
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.
Comments9 pages, 4 figures, 5 tables