arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39785cs.CV

像人类一样观察:从运动学习无监督分割一切

Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision

  • AI Research(360人工智能研究院)
  • University of Ottawa(渥太华大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Weijian Jian, Xiaoyue Zhang, Bin Xiao, Chunyu Xie, Yixiao He, Yutao Liu, Dawei Leng, Yuhui Yin

AI总结:

提出MoSA,一种利用大规模未标注视频中的运动信息学习可迁移对象先验的无监督分割框架,通过多粒度伪标签、对比学习和提示引导迁移,实现与全监督SAM相当的性能。

AI中文摘要:

Segment Anything Model(SAM)严重依赖大规模人工标注,这成为模型扩展的根本瓶颈。虽然无监督方法尝试从运动中学到对象概念,但它们通常对移动实体过拟合,缺乏多粒度理解以及泛化到静态对象的能力。为了克服这些问题,我们提出了Motion-Grounded Segment Anything(MoSA),一个高度可扩展的无监督框架,从未标注视频中学习可迁移的对象性先验。MoSA分三个阶段逐步进行:(1)从大规模视频数据中自动生成多粒度运动伪标签;(2)通过对比学习训练感知分组模型(PGM),以内化一种通用的、外观驱动的对象概念;(3)将学到的先验迁移到提示引导架构中,用于图像上的类SAM推理。在七个具有挑战性的基准(如COCO和ADE20K)上进行的大量零样本评估表明,MoSA显著优于现有无监督方法。值得注意的是,尽管使用了零人工标注,MoSA实现了与完全监督的SAM相当的分割性能。我们的发现表明,利用大规模未标注运动是替代基于标注的分割一切流程的一种可行且高度可扩展的方案。

英文摘要:

The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised methods attempt to learn object concepts from motion, they typically overfit to moving entities, lacking both multi-granularity understanding and the ability to generalize to static objects. To overcome this, we introduce Motion-Grounded Segment Anything (MoSA), a highly scalable unsupervised framework that learns a transferable objectness prior from unlabeled videos. MoSA operates in three progressive stages: (1) automatically generating multi-granularity motion pseudo-labels from large-scale video data; (2) training a Perceptual Grouping Model (PGM) via contrastive learning to internalize a generalized, appearance-driven concept of objects; and (3) transferring this learned prior into a prompt-guided architecture for segment-anything-style inference on images. Extensive zero-shot evaluations across seven challenging benchmarks (e.g., COCO and ADE20K) demonstrate that MoSA significantly outperforms existing unsupervised methods. Notably, despite using zero manual annotations, MoSA achieves segmentation performance comparable to the fully supervised SAM. Our findings reveal that harnessing large-scale unlabeled motion is a feasible and highly scalable alternative to annotation-driven segment-anything pipelines.

补充信息

↑