发表机构
School of Remote Sensing and Geomatics Engineering, Nanjing University of Information Science and Technology; Nanjing Center, China Geological Survey; School of Automation, Southeast University(南京信息工程大学遥感与测绘工程学院; 中国地质调查局南京中心; 东南大学自动化学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对遥感图像开放词汇语义分割中SAM3对几何变换响应不一致的问题,提出基于二面体群变换的特征自适应流形修复方法,通过视图选择、锚定修复和测试时自适应,在八个基准上平均mIoU达55.6%。
AI 中文摘要
遥感图像中物体的语义信息通常对二面体群D4的几何变换具有不变性。然而,基于SAM3的开放词汇语义分割(OVSS)方法往往对不同几何变换表现出不一致的响应。为利用这一特性并提高遥感图像OVSS的稳定性,我们提出了一种基于二面体群几何变换的特征自适应流形修复方法。首先,我们引入了多尺度谐波引导的D4视图选择(MH-D4VS),从一组几何变换视图中选取互补的候选视图。其次,我们提出了原始视图锚定的自适应流形修复(OAMR),以原始视图为锚点,利用可靠的跨视图信息选择性地修复局部不可靠的视觉特征。最后,我们为SAM3开发了像素解码器测试时自适应(PD-TTA),在推理过程中仅在线微调GroupNorm层的参数,从而增强模型适应样本级分布偏移的能力。实验结果表明,所提方法在八个遥感语义分割基准上实现了平均mIoU为55.6%,并在不同的基于SAM3的推理框架下均带来一致的性能提升。
英文摘要
The semantic information of objects in remote sensing images is typically invariant to geometric transformations from the dihedral group D4. However, SAM3-based open-vocabulary semantic segmentation (OVSS) methods often exhibit inconsistent responses to different geometric transformations. To exploit this property and improve the stability of OVSS for remote sensing images, we propose a feature-adaptive manifold repair method based on dihedral-group geometric transformations. First, we introduce multi-scale harmonic-guided D4 view selection (MH-D4VS) to select complementary candidate views from a set of geometrically transformed views. Next, we propose original-view-anchored adaptive manifold repair (OAMR), which uses the original view as an anchor and reliable cross-view information to selectively repair locally unreliable visual features. Finally, we develop pixel decoder test-time adaptation (PD-TTA) for SAM3, which fine-tunes only the parameters of the GroupNorm layers online during inference, thereby enhancing the model's ability to adapt to sample-level distribution shifts. Experimental results show that the proposed method achieves an average mIoU of 55.6% across eight remote sensing semantic segmentation benchmarks and delivers consistent performance improvements under different SAM3-based inference frameworks.