S4VY:前馈4D视觉几何中的任意分割
S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
- Texas A&M University(德克萨斯A&M大学)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
S4VY提出基于前馈4D视觉几何的Segment Anything模型,通过时空查询解码生成类无关4D实例掩码,并利用智能体工具实现自然语言定位,在静态和动态场景中达到最先进性能。
AI中文摘要:
动态场景中的精确实例分割对于机器人和自动驾驶等下游应用至关重要。现有的Segment Anything模型主要基于2D图像或视频掩码进行操作,并通过顺序记忆保持身份一致性,而基于前馈视觉几何构建的可提示4D实例分割仍未得到充分探索。我们提出了S4VY,一种基于前馈4D视觉几何构建的Segment Anything模型。从一组RGB观测中,S4VY通过时空查询解码器将共享的视觉几何特征转换为一组全面的类无关4D实例掩码,每个持久对象查询在所有观测中绑定一个实体。该表示支持与提示无关的分割以及基于点和框的条件选择,无需种子掩码或时间排序。我们进一步开发了一个智能体工具,用于在4D场景的大观测空间中进行自然语言定位。主动树搜索在不扫描每个固定窗口的情况下识别相关帧;双流定位器通过互补的边界框预测和对象查询匹配,将细粒度的VLM视觉先验与几何一致的实例特征相结合;独立的评判器从其预测中选择最终的4D实例掩码。大量实验表明,在涵盖静态和动态场景的统一评估下,S4VY实现了最先进的4D实例分割和强大的语言引导定位性能。
英文摘要:
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.