迈向全面的3D定位:通过视觉语言模型进行方向定位
Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文提出方向定位任务,构建ReferOri数据集和OG-VLM模型,通过6D方向和轴向对称预测,显著提升3D场景中对象方向理解能力。
中文摘要 AI 辅助
定位是空间视觉语言模型的核心能力,然而现有工作大多仅关注所指对象的位置。许多3D任务还需要知道对象的方向。尽管现有的3D VLM可能预测有方向的边界框,但框的姿态并不能明确捕捉对象中心的方向或对称性引起的歧义。我们引入了方向定位,这是一种指代定位任务,从单视图或多视图场景中的语言或框查询预测对象的6D方向和轴向对称性。为支持此任务,我们构建了ReferOri,包含331K多视图和387K单视图方向定位查询,这些查询通过可扩展的重建、一致性检查和人工验证获得。我们进一步提出了OG-VLM,它通过结构化的框/方向输出、符号和对称性标记以及几何感知的辅助损失来调整3D VLM。在单视图和多视图基准测试中,OG-VLM显著优于具有方向感知的VLM基线,并在场景级指代基准上超越了对象级方向基础模型,表明显式方向定位是一种区别于定位的可学习能力。下游结果验证了其对方向相关空间推理的益处。
英文摘要
Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-induced ambiguities. We introduce orientation grounding, a referring grounding task that predicts an object's 6D orientation and axial symmetry from a language or box query in single-view or multi-view scenes. To support this task, we construct ReferOri, with 331K multi-view and 387K single-view orientation-grounding queries obtained through scalable reconstruction, consistency checking, and human verification. We further present OG-VLM, which adapts a 3D VLM with structured box/orientation outputs, sign and symmetry tokens, and geometry-aware auxiliary losses. Across single-view and multi-view benchmarks, OG-VLM substantially outperforms orientation-aware VLM baselines and surpasses object-level orientation foundation models on scene-level referring benchmarks, showing that explicit orientation grounding is a distinct and learnable capability beyond localization. Downstream results validate its benefit for orientation-related spatial reasoning.