发表机构
Yonsei University; NVIDIA(延世大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RelationVGGT提出一种前馈、无姿态的多视角三维空间关系分割框架,结合视觉与几何基础模型及关系变换器,无需类别名称和逐场景优化,并基于ScanNet++提供自动化标注流程。
AI 中文摘要
近期三维重建的进展已从逐场景优化发展到前馈推理,语义场景理解也随之进步——然而现有方法仍局限于以物体为中心的感知,忽视了物体之间的空间关系。我们在前馈、无姿态的多视角设置中提出了三维空间关系分割任务:给定一个视觉指定的主体和一个关系文本查询,模型在不接收其类别名称的情况下跨视角分割目标。为此,我们提出了RelationVGGT,一种新颖的前馈框架,它将视觉基础模型的语义特征与三维几何基础模型的几何感知表示相结合,并利用关系变换器进行主体条件化的跨视角关系预测——既不需要逐场景优化,也不需要已知相机姿态。我们还提供了一个基于ScanNet++的完全自动化标注流程,利用VLM和LLM,为该新任务生成可扩展的训练数据。
英文摘要
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
Comments10 pages. Accepted to NeurIPS 2026 (poster). Project page: https://relationvggt.github.io/