arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GoDeep:通过语言空间提升实现免标注的开放词汇3D场景理解

GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting

Thodoris Betsas, Anastasios Doulamis, Andreas Georgopoulos

arXiv 2609.09082首次发表:更新:

发表机构

School of Rural, Surveying and Geoinformatics Engineering, NTUA(希腊国家技术大学农村、测量与地理信息工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GoDeep利用视觉-语言模型作为翻译器,在纯语言嵌入空间中直接聚合实体级描述,实现免标注的开放词汇3D场景理解,无需3D训练数据,性能优于CLIP基线并支持点级可解释性。

AI 中文摘要

开放词汇3D语义分割方法通常将CLIP特征提升到3D中。这会将点嵌入到一个联合视觉-语言空间中,而该空间在组合任务上的表现已知类似于词袋模型。此外,即使是免标注的变体,也通常需要大型3D训练语料库和针对每个领域的专用3D编码器。相反,我们仅将视觉-语言模型用作翻译器。它为每个带姿态的图像生成结构化的、实体级别的描述。这些描述直接在通用的、仅限语言的嵌入空间中进行接地、投影和聚合,无需3D训练语料库或编码器。在ScanNet++上,我们的流程与在ScanNet上训练的强免标注基线相当。在一个包含5栋建筑的文化遗产基准上,原始分数最初偏向于基于CLIP的变体,但一次系统的词汇校正逆转了这一排名。这一效果通过在不同类别上的第二次独立校正得到证实,表明语言空间嵌入更忠实地跟踪物理内容。这种保真度扩展到ScanNet++上真正超出词汇表(OOV)的物体,证明语言空间嵌入在区分存在与不存在物体方面远比基于CLIP的嵌入更清晰。GoDeep还能在场景中定位这些OOV物体,且无需任何2D-3D标注。由于所有表示保持为离散文本,预测在点级别也是可解释的。最后,利用一种有利于精确而非仅频繁观察的启发式加权以及GoDeep的可解释性,我们提出了一种聚合策略,作为概念验证,该策略有利于更精细的元素定位。

英文摘要

Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a joint vision-language space known to behave like a bag-of-words on compositional tasks. Furthermore, even annotation free variants often require a large 3D training corpus and a dedicated 3D encoder per domain. Instead we use a vision-language model purely as a translator. It produces structured, entity-level descriptions of each posed image. These descriptions are grounded, projected, and aggregated directly in a general-purpose, language-only embedding space, with no 3D training corpus or encoder required. On ScanNet++, our pipeline is competitive with strong annotation free baselines trained on ScanNet. On a 5-building cultural heritage benchmark, raw scores initially favor a CLIP-based variant, but a single systematic vocabulary correction reverses this ranking. An effect confirmed by a second, independent correction on a different class, indicating that language-space embeddings track physical content more faithfully. This fidelity extends to genuinely out-of-vocabulary (OOV) objects on ScanNet++ proving that language-space embeddings separate presence from absence objects far more sharply than CLIP-based embeddings do. GoDeep also localize these OOV objects within the scene, all without any 2D-3D annotation. Because every representation remains discrete text, predictions are also explainable at the point level. Finally, exploiting both a heuristic weighting, that favors precise over merely frequent observations and GoDeep's explainability property, we propose an aggregation strategy, as a proof of concept, that favors finer elements localization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑