arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token聚类与语义序列Mamba用于高光谱图像分类

Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu

arXiv 2609.28580首次发表:更新:

发表机构

University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高光谱图像分类中光谱-空间异质性问题,提出Token聚类与语义序列Mamba(STMamba),通过密度感知聚类和四叉树动态选择构建语义连贯序列,并利用并行空间-光谱Mamba模块捕获长距离依赖,在三个大规模基准数据集上超越现有最先进方法。

AI 中文摘要

尽管高光谱图像(HSIs)提供了丰富的光谱-空间信息,但由于光谱-空间异质性和复杂的空间结构,精确的像素级分类仍然具有挑战性。现有的视觉状态空间模型(Mamba)通常根据预定义的空间邻域构建序列,而没有明确考虑语义相似性或空间非平稳性。为了解决这一局限性,我们提出了Token聚类与语义序列Mamba(STMamba),该方法将稀疏token组织成语义连贯的序列,用于高光谱图像分类,具有以下特点。首先,在宏观层面,分层编码器-解码器通过Token聚类模块(TCM)逐步选择语义token,并使用无参数的跨尺度邻域注意力(CNA)上采样器恢复密集特征。其次,在微观层面,TCM首先通过密度感知聚类识别代表性的聚类中心,并基于特征相似性估计软隶属度。然后,基于四叉树的动态选择策略从每个语义聚类中保留稀疏且空间分布的token,形成连贯的语义token序列,同时减少冗余的像素级表示。第三,并行的空间和光谱语义序列Mamba(SWSM)模块在同质语义token序列内捕获互补的长距离空间和光谱依赖性,同时抑制异质区域之间的无关交互。在三个大规模基准数据集上的实验结果表明,STMamba在定量和定性结果方面均优于最先进(SOTA)方法。

英文摘要

Although hyperspectral images (HSIs) provide rich spectral-spatial information, accurate pixel-level classification remains challenging because of spectral-spatial heterogeneity and complex spatial structures. Existing vision state space models (Mamba) typically construct sequences according to predefined spatial neighborhoods, without explicitly accounting for semantic similarity or spatial non-stationarity. To address this limitation, we propose Token Clustering and Semantic Sequence Mamba (STMamba), which organizes sparse tokens into semantically coherent sequences for hyperspectral image classification with the following features. First, at the macro level, a hierarchical encoder decoder progressively selects semantic tokens with the Token Clustering Module (TCM) and restores dense features using a parameter-free Cross-scale Neighborhood Attention (CNA) Upsampler. Second, at the micro level, TCM first identifies representative cluster centers through density-aware clustering and estimates soft memberships based on feature similarity. A quadtree-based dynamic selection strategy then retains sparse and spatially distributed tokens from each semantic cluster, forming coherent semantic-token sequences while reducing redundant pixel-wise representations. Third, parallel Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture complementary long-range spatial and spectral dependencies within homogeneous semantic token sequences while suppressing irrelevant interactions across heterogeneous regions. Experimental results on three large-scale benchmark datasets demonstrate that STMamba outperforms the SOTA methods with respect to quantitative and qualitative results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑