AI 中文总结
针对遥感语义分割标注成本高的问题,提出无需训练的DinoSplat-OV框架,结合DINOv3的文本感知拉普拉斯传播与高斯溅射上采样模块,在多数据集上取得优于现有方法的性能,填补了DINO系列模型在该领域的空白。
AI 中文摘要
遥感语义分割受限于像素级标注的高昂成本,这推动了无需训练的开放词汇方法的发展。近期发布的DINOv3为独立DINO骨干网络配备了图像-文本对比学习,从而为开放词汇分割开辟了可能。我们提出了DinoSplat-OV,这是一种无需训练的框架,可适配DINOv3用于遥感任务,无需微调或额外预训练。针对遥感影像的密集分布、多尺度特性及大尺寸特点,我们设计了两个核心模块:其文本感知拉普拉斯传播模块通过结合文本语义亲和性与局部视觉相似性对块级预测进行去噪,提升区域一致性的同时保留边界;其高斯溅射上采样模块通过RGB引导的各向异性聚合与测试时优化重建像素级特征。全局锚定滑动窗口策略进一步支持大尺寸影像。在UDD5、DOTA和LoveDA数据集上的实验表明,该方法相较于现有无需训练的方法具有可比或更优的性能,有效填补了DINO系列模型在无需训练的开放词汇分割领域的空白,为该方向的进一步发展提供了可行的新路径。
英文摘要
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.