arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从上方进行可提示的概念分割:评估SAM 3在遥感中的零样本和单样本能力

Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing

Mohammad Dabaja, Turgay Celik

arXiv 2607.09583首次发表:更新:

发表机构

University of Agder(阿格德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在严格零样本和单样本约束下,对SAM 3在遥感场景分类等任务中的能力进行评估。核心方法是对SAM 3结构调整并隔离提示模态来诊断对齐机制,还制定代理评估协议。主要贡献是揭示跨模态干扰,表明SAM 3避免过拟合但受分辨率等限制,指明微调方向。

AI 中文摘要

大规模基础模型的部署,如Segment Anything Model 3(SAM 3),有望向开放词汇、无需训练的计算机视觉转变。然而,其在分布外对地球观测图像复杂的自上而下几何结构的泛化能力仍未得到充分量化。受SAM 3在高度专业化领域性能差异的驱动,我们在严格的零样本和单样本约束下,对遥感场景分类、目标检测和实例分割进行了全面的多任务实证评估。为此,我们对SAM 3进行了结构调整,将其解耦的二进制存在头重新用作独立的零样本分类器。此外,通过系统地隔离五种配置中的文本和视觉提示模态,我们明确诊断了模型多模态解码器内的对齐机制。我们的发现揭示了严重的跨模态干扰:视觉提示成功地将解码器与复杂的遥感几何结构对齐,而文本提示则注入了未对齐且处于地面水平的语义偏差,从而严重降低了坐标回归。为了在无需资源密集型训练的情况下对这些能力进行基准测试,我们为广义零样本任务(场景分类和实例分割)制定了一种新颖的无需训练的代理评估协议。最终,我们的结果表明,SAM 3避免了传统领域适应模型中常见的过拟合问题,在分割任务中获得了较高的调和平均分数。然而,它仍然受到亚像素分辨率限制和开销语义盲点的根本约束,为其多模态解码器的参数高效地理空间微调指明了明确的方向。

英文摘要

The deployment of large-scale foundation models, such as the Segment Anything Model 3 (SAM 3), promises a transition toward open-vocabulary, training-free computer vision. However, their capacity to generalize out-of-distribution to the complex, top-down geometric structures of Earth Observation imagery remains largely unquantified. Driven by SAM 3's performance disparities in highly specialized domains, we present a comprehensive, multi-task empirical evaluation across remote sensing scene classification, object detection, and instance segmentation under strict zero-shot and one-shot constraints. To achieve this, we introduce a structural adaptation of SAM 3 by repurposing its decoupled binary presence head into a standalone zero-shot classifier. Furthermore, by systematically isolating textual and visual prompt modalities across five configurations, we explicitly diagnose the alignment mechanics within the model's multimodal decoder. Our findings reveal severe cross-modal interference: while visual prompts successfully align the decoder to complex remote sensing geometry, textual prompts inject misaligned, ground-level semantic bias, actively degrading coordinate regression. To benchmark these capabilities without resource-intensive training, we formulate a novel training-free proxy evaluation protocol for Generalized Zero-Shot tasks (scene classification and instance segmentation). Ultimately, our results demonstrate that SAM 3 avoids the overfitting commonly seen in legacy domain-adapted models, achieving high Harmonic Mean scores in segmentation tasks. However, it remains fundamentally constrained by sub-pixel resolution limits and overhead semantic blind spots, charting a definitive mandate for parameter-efficient geospatial fine-tuning of its multimodal decoder.

Comments14 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑