arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15647cs.CVcs.AI

面向超高分辨率(VHR)遥感图像分割的分层自适应特征细化网络

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin

AI总结:

针对VHR遥感图像分割中预训练分层编码器利用难的问题,提出HAFR-Net框架,含HG-SAF、FRA、CATP模块,在多个数据集上较UPerNet基线取得显著mIoU提升。

AI中文摘要:

超高分辨率(VHR)遥感图像的语义分割越来越受益于强大的预训练分层编码器,但利用其多阶段表示仍然存在困难:邻近区域对精细细节和语义上下文的平衡需求不同,激进的特定任务变换会扰动有用的预训练特征,而传统的语义监督提供的结构指导有限。本文提出HAFR-Net,一种渐进式细化框架,该框架自适应组织并保守细化分层表示,而非用单一解码器变换替换它们。异质性引导的阶段自适应融合(HG-SAF)基于局部特征变化预测密集阶段权重;频率残差适配器(FRA)随后通过有界的零初始化残差分支注入频率信息,该分支将融合表示作为参考;混淆感知三先验解码器(CATP)最终利用边界、目标性和训练衍生的类关系线索对预测进行正则化。在匹配的Swin-B训练和单尺度推理协议下,HAFR-Net在ISPRS Vaihingen、ISPRS Potsdam、LoveDA和OpenEarthMap上分别达到84.12%、87.86%、55.17%和67.70%的mIoU,较匹配的UPerNet基线分别提升0.55、0.95、1.55和1.84个百分点。控制分析进一步表明,该方法实现了超越仅内容路由的一致空间重加权,较匹配的空间和光谱替代方案提升了边界与精细结构的准确率,并减少了预声明类对的混淆。

英文摘要:

Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.

补充信息

↑