发表机构
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University; College of Computer Science, Nankai University; Tsinghua Shenzhen International Graduate School, Tsinghua University; Department of Electronic Engineering, Tsinghua University(上海交通大学信息科学与工程学院; 南开大学计算机学院; 清华大学深圳国际研究生院; 清华大学电子工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对遥感图像指称分割的精度与效率平衡问题,提出DiCoR框架,通过解耦指称消歧与轮廓重校准,在三个基准上实现最优精度且效率优于代表性方法。
AI 中文摘要
遥感图像指称分割(Referring Remote Sensing Image Segmentation, RRSIS)旨在勾勒出遥感图像中由自然语言表达式指定的目标。现有方法主要遵循联合融合分割(Joint Fusion Segmentation, JFS)或解耦提示分割(Decoupled Prompt Segmentation, DPS)范式。JFS效率较高,但常因指称定位与掩码勾勒在统一目标下优化而精度有限;DPS则利用空间提示和基础分割器将定位与掩码生成分离,代价是更高的内存消耗和推理延迟。为弥合该差距,本文提出DiCoR,一种基于高效JFS流水线构建的解耦指称消歧与轮廓重校准框架。DiCoR解决两大关键挑战:从歧义候选中区分正确指称,以及在定位后细化粗糙掩码。一种消歧感知的定位引导策略利用自适应语言线索对显著候选区域排序,并将所得定位先验注入融合特征;轻量轮廓重校准模块在定位轮廓监督下进一步预测粗糙逻辑的残差修正,以有限计算开销提升掩码质量。在RefSegRS、RRSIS-D和RISBench三个基准上的实验表明,DiCoR在所有基准中取得最优分割精度;在RefSegRS上,其较具有竞争力的JFS方法将mIoU和gIoU分别提升5.28%和2.87%,同时比代表性DPS方法运行速度快4.7%,展现出良好的精度-效率权衡。代码可在this https URL获取。
英文摘要
Referring remote sensing image segmentation (RRSIS) aims to delineate targets specified by natural language expressions in remote sensing imagery. Existing methods mainly follow joint fusion segmentation (JFS) or decoupled prompt segmentation (DPS). JFS is efficient but often suffers from limited accuracy because referent localization and mask delineation are optimized under a unified objective, whereas DPS separates localization from mask generation using spatial prompts and foundation segmenters at the cost of higher memory consumption and inference latency. To bridge this gap, we propose DiCoR, a decoupled referent disambiguation and contour recalibration framework built on an efficient JFS pipeline. DiCoR addresses two key challenges: distinguishing the correct referent from ambiguous candidates and refining coarse masks after localization. A disambiguation-aware localization guidance strategy ranks salient candidate regions with adaptive linguistic cues and injects the resulting localization prior into fused features. A lightweight contour recalibration module further predicts residual corrections to coarse logits under localized contour supervision, improving mask quality with limited computational overhead. Experiments on RefSegRS, RRSIS-D, and RISBench show that DiCoR achieves the best segmentation accuracy across all three benchmarks. On RefSegRS, it improves mIoU and gIoU by 4.90% and 2.16% over a competitive JFS method while running 2.4-6.5 times faster than the evaluated DPS method, demonstrating a favorable accuracy-efficiency trade-off. Code is available at https://github.com/zyGao1126/DiCoR.
CommentsRevised version with additional experiments and updated results