AI 中文总结
针对遥感指称分割中架构弱耦合与语义偏差问题,本文提出CROSS范式,通过LGCD与PSCL实现性能突破,成为RRSIS的稳健新方案。
AI 中文摘要
遥感图像指称分割(RRSIS)通过结合视觉语言模型(VLM)与分割任何模型(SAM)已取得显著进展,但该进展在很大程度上依赖强大的预训练能力,同时未充分解决两个基本局限:(1)架构弱耦合,单向流迫使模型依赖粗粒度VLM提示,且浪费SAM的像素级结构指导,导致定位漂移;(2)以对象为中心的语义偏差,模型过度强调主导对象语义,对RRSIS至关重要的空间推理不敏感。基于这些观察,本文提出用于RRSIS的紧密集成范式CROSS。首先,引入语言引导级联蒸馏(LGCD)以弥合架构差距,将SAM的几何亲和力作为软正则化项蒸馏到VLM中间层,注入密集结构先验以优化定位。其次,视角空间对比学习(PSCL)通过挖掘掩码过滤的欺骗性干扰项和空间语言反事实作为硬负样本,施加跨锚定约束,明确打破语义捷径以确保真正的逻辑一致性。在RRSIS基准上的大量实验表明,CROSS达到了最先进的性能,即使在严重的空间描述扰动下也能保持精确的定位,是RRSIS的一种稳健新范式。
英文摘要
Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
CommentsAccepted at the European Conference on Computer Vision (ECCV) 2026. 20 pages, 6 figures, and 5 tables. Tingzhang Luo and Ruizhong Liu contributed equally. Jianyuan Guo is the corresponding author. Project page: https://clarence-cv.github.io/CROSS/