arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03763cs.CV

TDVR:用于零样本3D视觉定位的联合文本消歧与视点推理框架

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

Qingxi Du, Junbo Wang, Yuke Li, Yining Zhu

AI总结:

本文提出TDVR框架,通过联合文本消歧与视点推理解决零样本3D视觉定位中的文本歧义与视点不足问题,在ScanRefer数据集上的Acc@0.25和Acc@0.5指标较现有最优方法分别提升15.25%和14.46%

AI中文摘要:

零样本3D视觉定位旨在基于文本描述和3D视觉输入定位特定对象,但现有方法的有效性受限于查询文本的歧义性和视点不足。为解决这些问题,本文提出TDVR,一个无需训练的推理框架,用于对输入文本进行消歧并为零样本3D视觉推理推断准确视点。首先,从3D点云中的检测实例构建语义3D场景图;随后,将原始查询、外观及空间关系描述输入LLM进行融合,以对初始输入进行消歧,并利用思维链推理生成消歧查询的结构化表示;接着,将场景图和结构化查询作为输入,通过视点推理获取最优视点,解决定位过程中的视点缺失问题;基于获取的最优视点,进一步区分干扰对象,使模型具备区分相似实例的能力;之后,通过计算特征向量的相似度,将类别文本和外观图像与查询进行匹配;最后,通过整合视点得分、混淆得分、类别得分和外观得分来识别目标对象。在公开ScanRefer数据集上的实验结果表明,与现有方法相比,TDVR在Acc@0.25和Acc@0.5上分别超出现有最优方法15.25%和14.46%,验证了其在解决查询文本歧义性和视点不足问题上的有效性。

英文摘要:

Zero-shot 3D visual grounding aims to localize specific objects based on textual descriptions and 3D visual input. However, the effectiveness of existing methods is significantly hindered by the ambiguous query text and deficient viewpoints. To address these issues, we propose TDVR, a training-free reasoning framework that disambiguates the input text and infers accurate viewpoints for zero-shot 3D visual grounding. First, we construct semantic 3D scene graph from the detected instances in the 3D point cloud. Subsequently, we put the original query, appearance and spatial relationship descriptions into the LLM for fusion, thereby disambiguating the initial input. We leverage chain-of-thought reasoning to generate the structured representation of disambiguated query. Then taking the scene graph and structured query as input, we get the optimal view via viewpoint reasoning to solve the problem of missing viewpoints during grounding. Based on the obtained optimal viewpoint, we further discriminate the distracting objects, enabling the model with the ability to distinguish similar instances. After that, we match the category text and appearance images with the query by computing the similarity of feature vectors. Finally, the target object was identified by integrating the viewpoint score, confusion score, category score, and appearance score. Compared with previous methods, our TDVR has stronger capabilities in viewpoint reasoning, similar object discrimination, and ambiguous query understanding. Experimental results on the public ScanRefer dataset show that our method outperforms the existing state-of-the-art methods by 15.25% and 14.46% in Acc@0.25 and Acc@0.5 respectively, demonstrating the effectiveness of our TDVR in addressing ambiguous query text and deficient viewpoints.

补充信息

↑