发表机构
School of Computer Science, China University of Geosciences; Systems Hub, The Hong Kong University of Science and Technology (Guangzhou); Department of Computer Science, City University of Hong Kong; School of Geodesy and Geomatics, Wuhan University; Data Science and AI Innovation Research Promotion Center, Shiga University(中国地质大学计算机学院; 香港科技大学(广州)系统中心; 香港城市大学计算机系; 武汉大学测绘学院; 滋贺大学数据科学与人工智能创新研究促进中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GeoSeg-OV,通过解耦辅助VFM特征与视觉-文本匹配,将其用作结构引导,结合SGA与CAD模块,在HRLC基准上优于当前最优方法,且具备跨域零样本泛化能力。
AI 中文摘要
开放词汇遥感分割是近年兴起的极具潜力的范式,可实现自然语言指定的任意类别的像素级识别,包括训练时未见过的类别。然而,由异质区域、空间分辨率和采集平台导致的地理空间域偏移会削弱视觉-文本匹配能力,限制跨数据集泛化性。近期研究尝试引入辅助视觉基础模型(VFMs),通常将其特征与文本嵌入结合作为额外匹配证据,但该策略可能引入不一致的匹配信号,且未充分利用VFMs对结构敏感的表征。为此,本文提出GeoSeg-OV,将辅助VFM特征与视觉-文本匹配解耦,重新将其用作代价聚合与解码的结构引导。GeoSeg-OV从多旋转CLIP特征构建方向鲁棒的代价体,同时冻结的VFM并行提取多尺度结构敏感特征。本文提出结构引导聚合(SGA),将代价标记与CLIP语义引导及VFM派生的成对结构偏差整合,实现连贯的空间传播,随后进行文本条件下的逐类推理;还提出代价感知解码(CAD),基于当前解码器上下文自适应地优化并融合多尺度语义与结构引导。在覆盖六大洲七个数据集的全球高分辨率土地覆盖(HRLC)基准上,GeoSeg-OV在两种训练设置下均以平均交并比(mIoU)分别超出当前最优方法2.5和2.7个百分点。大规模零样本案例研究进一步证明,该方法无需目标域标注或重新训练即可跨地理域和类别系统泛化。
英文摘要
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization. Recent attempts have begun to incorporate auxiliary vision foundation models (VFMs), typically coupling their features with text embeddings as additional matching evidence. However, this strategy may introduce inconsistent matching signals while leaving the structure-sensitive representations of VFMs insufficiently exploited. We therefore propose GeoSeg-OV, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding. GeoSeg-OV constructs an orientation-robust cost volume from multi-rotation CLIP features, while a frozen VFM extracts multi-scale structure-sensitive features in parallel. We propose Structure-Guided Aggregation (SGA), which integrates cost tokens and CLIP semantic guidance with VFM-derived pairwise structural biases for coherent spatial propagation, followed by text-conditioned class-wise reasoning. We further introduce Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context. On the global High-Resolution Land Cover (HRLC) benchmark spanning seven datasets across six continents, GeoSeg-OV outperforms the state-of-the-art by +2.5 and +2.7 average mIoU under two training settings. A large-scale zero-shot case study further demonstrates its generalization across geographic domains and category systems without target-domain annotations or retraining.
CommentsCode and benchmark: https://github.com/zzaiyan/GeoSeg-OV