HyperSAM:用于高光谱遥感的可提示基础模型
HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing
- Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
- Helmholtz-Zentrum Dresden-Rossendorf(亥姆霍兹德累斯顿罗森多夫研究中心)
- University of Iceland(冰岛大学)
- Griffith University(格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HyperSAM提出结合数据合成与SAM3架构的高光谱基础模型,通过物理信息生成和伪标签训练,在多种任务上实现强泛化,证明高质量合成数据优于噪声监督。
AI中文摘要:
高光谱遥感提供了密集的光谱测量,这对于物质级地球观测不可或缺,然而构建通用的高光谱基础模型仍然困难重重。有两个瓶颈尤其具有限制性。首先,大型高光谱语料库很少能同时提供高空间分辨率和可靠的密集标注。其次,许多高光谱模型仍然几乎从零开始训练,因此现代视觉基础模型学到的几何和交互先验没有被充分利用。为了缓解这些问题,我们提出了HyperSAM,一个可提示的高光谱基础模型,它将数据驱动的高光谱合成流程与基于Segment Anything Model 3(SAM3)的光谱适应架构相结合。在数据方面,HyperSAM通过物理信息丰度转移生成器,从高分辨率SpaceNet多光谱图像中合成全光谱高光谱立方体,而SAM3派生的伪掩码则提供对象中心监督。在模型方面,最新实现使用冻结的SAM3 RGB图像分支、从RGB视觉变换器(ViT)初始化的可训练高光谱侧编码器、ControlNet风格的零初始化特征注入,以及轻量级混合专家掩码精化器。为了增强对噪声伪标签的训练鲁棒性,引入了Cross-modal Sample Selection(CromSS)风格的置信度选择用于噪声标签加权。大量实验表明,HyperSAM在各种高光谱任务(例如分类、异常检测、变化检测、目标检测和航空溢油测绘)上获得了强大的泛化能力,并且高质量的合成高光谱数据可能比简单扩展噪声高光谱监督更有效。
英文摘要:
Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.