arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Lang3DSeg:基于点 Transformer 的无标注开放词汇 3D 分割

Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

Cigdem Kokenoz, Amir Salarpour, Alkim Domeke, Christopher Salas, Pedram MohajerAnsari, Long Cheng, Mert D. Pesé, Bing Li

arXiv 2610.00855首次发表:更新:

发表机构

Clemson University(克莱姆森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Lang3DSeg 首次将点 Transformer 用于室外 LiDAR 无标注开放词汇分割,通过类别优先级合成掩码与深度截断纠正投影噪声,在 nuScenes 和 SemanticKITTI 上取得最高 mIoU。

AI 中文摘要

精确的 3D 语义感知对于安全的自主导航至关重要。然而,有监督的 LiDAR 分割仍然受限于封闭的分类体系以及逐点手动标注的高昂成本。开放词汇方法通过将 2D 视觉-语言模型的输出投影到 LiDAR 数据上,并将其蒸馏到 3D 网络中,从而避免了这一成本。这些方法几乎完全依赖于基于体素(voxel)的稀疏卷积,而点 Transformer 迄今为止仅限于室内环境,在室内环境中 3D 数据密集且有界。我们提出了 Lang3DSeg,它将点 Transformer 确立为无标注开放词汇室外 3D LiDAR 分割的骨干网络,并且无需几何预训练即可从头训练。这种训练范式需要解决 2D 到 3D 标签投影中固有的噪声问题;具体而言,朴素投影常常遭受深度模糊(depth ambiguity)的影响,即位于物体后面的点被错误地赋予该物体的语义标签。因此,我们使用显式的类别优先级规则合成掩码,并在每个投影实例的深度分布的第一个间隙处截断该实例,直接纠正投影误差,而不是在配准序列上对其进行平均。Lang3DSeg 在 nuScenes 验证集上达到了 52.8% 的 mIoU,在 SemanticKITTI 上达到了 41.4%,在这两个基准上均为已发表的无标注方法中的最高水平。每个 3D 语义分割均基于单次 LiDAR 扫描,并且推理无需运行视觉-语言模型即可实时进行。

英文摘要

Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.

Comments9 pages, 3 figures, 4 tables. Submitted to IEEE ICRA 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑