发表机构
University of Illinois at Urbana-Champaign; Meta(伊利诺伊大学厄巴纳-香槟分校; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SAM-V将几何模型VGGT特征集成到SAM中,通过提示融合机制实现端到端的多视图实例分割,无需离线掩码匹配或三维重建,在IGGT基准上显著提升IoU和召回率。
AI 中文摘要
一致的多视图物体分割对于三维感知和机器人技术至关重要,但在严重的视角变化和遮挡变化下仍然具有挑战性。现有方法通常在点云上执行三维实例分割,或依赖离线的二维掩码匹配流程。然而,三维实例分割受限于稀缺的三维标注,而离线二维匹配则面临跨帧物体身份模糊的问题。为了联合利用强大的二维和三维先验,我们提出了SAM-V(几何感知的任意分割模型用于多视图实例分割)。SAM-V不是通过事后匹配来组合这两种先验,而是直接将前馈几何模型(VGGT)的特征集成到二维分割基础模型(SAM)中,并以端到端方式训练用于跨视图实例预测。SAM-V引入了一种提示融合机制,该机制用视图特定的相机令牌和局部VGGT特征丰富稀疏的SAM提示令牌,使提示表示既具有视图感知性又具有空间基础性,同时配备一个掩码解码器,该解码器关注密集的二维和三维特征。通过直接基于多视图几何条件化掩码解码,SAM-V在单次前向传播中即可生成提示对象的一致多视图分割,无需离线掩码匹配或显式三维重建。在IGGT 3D跟踪基准上,其中跨帧的一致实例身份直接决定性能,SAM-V在ScanNet++分割上相比最先进的多视图实例分割基线,将整体IoU提高了5个百分点,帧级召回率提高了12个百分点,并在零样本ScanNet分割的所有指标上领先。我们的代码和预训练模型可在该https URL获取。
英文摘要
Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance segmentation is limited by scarce 3D annotations, while offline 2D matching suffers from object identity ambiguity across frames. To leverage strong 2D and 3D priors jointly, we propose SAM-V (Geometry-Aware Segment Anything for Multi-View Instance Segmentation). Instead of combining the two priors through post-hoc matching, SAM-V directly integrates features from a feed-forward geometry model (VGGT) into a 2D segmentation foundation model (SAM), trained end-to-end for cross-view instance prediction. SAM-V introduces a prompt-fusion mechanism that enriches sparse SAM prompt tokens with view-specific camera tokens and local VGGT features, making the prompt representation both view-aware and spatially grounded, together with a mask decoder that attends to dense 2D and 3D features. By conditioning the mask decoding directly on multi-view geometry, SAM-V produces consistent multi-view segmentation of a prompted object in a single forward pass without offline mask matching or explicit 3D reconstruction. On the IGGT 3D tracking benchmark, where consistent instance identity across frames directly determines performance, SAM-V improves overall IoU by 5 points and frame-level recall by 12 points on the ScanNet++ split over the state-of-the-art multi-view instance segmentation baseline and leads on all metrics in the zero-shot ScanNet split. Our code and pretrained models are available at https://github.com/gong208/SAM-V.git.