arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VPRef:用于参考遥感图像分割的跨域基准

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang

arXiv 2609.16486首次发表:更新:

发表机构

James Cook University; University of Wisconsin-Madison; La Trobe University(詹姆斯·库克大学; 威斯康星大学麦迪逊分校; 拉筹伯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对参考遥感图像分割的跨域性能下降问题,提出VPRef基准数据集和基于SAM3与LoRA的域适应框架,仅修改1.08%参数即实现优越跨域分割。

AI 中文摘要

视觉-语言模型的快速发展将参考遥感图像分割(RRSIS)推至地球观测的前沿。然而,在实际部署中,在一种耦合的双漂移范式下,性能会严重下降:视觉域漂移源于跨空间分辨率不匹配和光谱变化,而文本逻辑漂移则源于不受约束、可变的用户输入粒度。为缓解这些瓶颈,本文建立了首个跨域RRSIS基准,命名为Vaihingen-Potsdam参考(VPRef)数据集,包含46,972个语言-图像-标注三元组,并组织成三层语言层级。在此基准之上,我们开发了一个定制的参数高效域适应基线,该基线基于Segment Anything Model(SAM3)并采用低秩适应(LoRA)技术。我们的框架通过伪标签驱动的自训练来抵消视觉分布差异,并通过随机多粒度文本提示混合来解决文本逻辑漂移。关键的是,跨消融变体的经验指标分布表明,跨模态语义鲁棒化与视觉域对齐之间可能存在解耦,证明语言变异性驱动细粒度语义不变性,而伪标签传播则控制宏观尺度的空间网格对齐。广泛的基准测试表明,所提出的框架在仅修改基础参数足迹的1.08%的情况下,实现了优越的跨域分割边界,为未来多模态遥感域适应研究建立了稳健的基线。数据集和代码将在该https URL上提供。

英文摘要

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches and spectral variations, alongside textual logic drift from unconstrained, variable user-input granularities. To mitigate these bottlenecks, this paper establishes the first cross-domain RRSIS benchmark, designated as the Vaihingen-Potsdam Referring (VPRef) dataset, comprising 46,972 language-image-annotation triplets organized into a three-tier linguistic hierarchy. Building upon this benchmark, we develop a tailored parameter-efficient domain adaptation baseline anchored on the Segment Anything Model (SAM3) via Low-Rank Adaptation (LoRA). Our framework counteracts visual distribution discrepancies through pseudo-label-driven self-training and addresses textual logic drift via random multi-granularity text prompt mixing. Crucially, the distribution of empirical metrics across ablative variants suggests a potential decoupling between cross-modal semantic robustification and visual domain alignment, demonstrating that linguistic variance drives fine-grained semantic invariance while pseudo-label propagation governs macro-scale spatial grid alignment. Extensive benchmarks demonstrate the proposed framework achieves superior cross-domain segmentation boundaries while modifying merely 1.08\% of the foundational parameter footprint, establishing a robust baseline for future multi-modal remote sensing domain adaptation research. The dataset and code will be available at https://github.com/quanweiliu/VPRef.

Comments12 pages, 7 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑