arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22834cs.CVcs.AI

SatOV:恢复空间先验用于遥感影像免训练开放词汇分割

SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery

Changhao Zhao, Linglin Zeng, Hai Liu

首次发表
浏览论文内容

中文总结 AI 辅助

SatOV通过ResQQ和SatUp在表示与分辨率两阶段恢复空间先验,实现免训练遥感开放词汇分割,并在多个基准上取得竞争性结果。

中文摘要 AI 辅助

开放词汇语义分割(OVS)在遥感影像中是一项具有挑战性的像素级任务,要求模型具备强大的泛化能力并适应遥感数据的空间特性。尽管现有的视觉-语言基础模型在通用领域表现良好,但其图像级分类设计削弱了高分辨率遥感分割所需的空间先验:深层特征变换过程中结构空间关系退化,下采样过程中细粒度空间细节丢失。为解决这些互补的缺陷,我们提出SatOV,一种用于开放词汇遥感分割的免训练框架,在表示流程的两个阶段恢复空间先验。具体而言,残差查询-查询注意力(ResQQ)从中间CLIP层提取查询-键自注意力,并通过残差组合与最终层的查询-查询注意力融合,恢复被最终层表示抑制的结构空间先验。空间调制上采样(SatUp)利用原始高分辨率RGB图像作为空间引导,结合空间特征调制与引导交叉注意力来重建像素级纹理和边界。在DOTA、UDD、LoveDA和Vaihingen上的大量实验表明,SatOV持续改进免训练OVS,并在定量和定性结果上达到与最先进方法竞争的水平。这些结果验证了在表示和空间分辨率阶段恢复空间先验对遥感开放词汇分割的有效性。

英文摘要

Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classification design weakens the spatial priors needed for high-resolution remote sensing segmentation: structural spatial relations are degraded during deep feature transformation, and fine-grained spatial details are lost during downsampling. To address these complementary deficiencies, we propose SatOV, a training-free framework for open-vocabulary remote sensing segmentation that restores spatial priors at two stages of the representation pipeline. Specifically, Residual QQ Attention (ResQQ) extracts Query-Key self-attention from an intermediate CLIP layer and fuses it with final-layer Query-Query attention via a residual combination, restoring structural spatial priors suppressed by the final-layer representation. Spatially Modulated Upsampling (SatUp) uses the original high-resolution RGB image as spatial guidance, combining spatial feature modulation with guided cross-attention to reconstruct pixel-level textures and boundaries. Extensive experiments on DOTA, UDD, LoveDA, and Vaihingen show that SatOV consistently improves training-free OVS and achieves competitive quantitative and qualitative results against state-of-the-art methods. These results validate the effectiveness of restoring spatial priors at both the representation and spatial-resolution stages for remote sensing open-vocabulary segmentation.

发表机构

  • College of Resources and Environment, Huazhong Agricultural University(华中农业大学资源与环境学院)

机构由 AI 辅助整理,请以论文原文为准。

↑