arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GuideGround:用于3D视觉定位的VLM引导语义理解与视点感知推理

GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

Yiwen Wang, Yuyang Deng, Yihao Long, Xi Zhao

arXiv 2608.00518首次发表:更新:

AI 中文总结

GuideGround是一种VLM引导的3D视觉定位框架,用VLM生成的语义描述替代闭集分类、逐视图保留假设并经VLM验证,在ReferIt3D基准上性能优于现有SOTA。

AI 中文摘要

3D视觉定位旨在从自然语言查询出发,在3D场景中定位目标物体,既需要细粒度的语义理解,也需要依赖视点的空间推理。现有方法通常将语义理解建模为辅助的闭集物体分类任务,并依赖多视图特征聚合进行视点推理,这限制了语义泛化能力,且削弱了视点特定证据。我们观察到,视觉语言模型(VLM)自然具备互补能力,包括开放词汇语义理解与全局场景感知。基于该见解,我们提出GuideGround,一种VLM引导的框架,它通过利用VLM进行语义增强和视点特定假设验证来补充而非替代特定任务的定位模型。具体而言,我们用VLM生成的物体语义描述替代辅助闭集物体分类,以增强语义理解;同时,我们不直接聚合多视图表示,而是通过逐视图定位保留视点特定的定位假设,并利用VLM在候选视点间对这些假设进行显式验证。在ReferIt3D基准上的大量实验表明,GuideGround始终优于先前的最先进方法;全面的 ablation 研究进一步证实了所提出的语义理解和视点推理策略的有效性。

英文摘要

3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification. Specifically, we replace auxiliary closed-set object classification with VLM-generated object semantic descriptions to enhance semantic understanding. Meanwhile, instead of directly aggregating multi-view representations, we preserve viewpoint-specific grounding hypotheses through per-view grounding and explicitly verify them using VLMs across candidate viewpoints. Extensive experiments on the ReferIt3D benchmark demonstrate that GuideGround consistently outperforms previous state-of-the-art methods. Comprehensive ablation studies further confirm the effectiveness of both the proposed semantic understanding and viewpoint reasoning strategies.

Comments18 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑