开放词汇三维目标检测与可提示分割
Open-vocabulary 3D object detection with promptable segmentation
- Institute of Data Science and Artificial Intelligence, Boğaziçi University(博阿齐奇大学数据科学与人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出利用可提示分割模型SAM3,以文本提示生成实例掩码并转换为三维框,实现无训练开放词汇三维目标检测,在nuScenes上达到0.298 mAP/0.348 NDS,并显著提升LiDAR检测器性能。
AI中文摘要:
自动驾驶的三维目标检测主要由基于大量人工标注三维框训练而成的检测器主导。这类检测器学习固定的类别列表,列表之外的一切均不可见。本文探讨能否以无训练和开放词汇的方式解决该任务。一个可提示分割模型(SAM3),以类别名称作为文本提示进行查询,在车辆的六个环视摄像头中提供实例掩码,并利用场景几何将这些掩码转换为度量三维框。核心是在nuScenes上进行的受控三阶段比较,其中二维检测保持固定,仅改变三维几何的来源。仅从图像预测的几何在官方协议下达到0.183的平均精度(mAP);使用无训练规则从同一掩码内的原始LiDAR点拟合框,在零标注成本下达到0.298 mAP / 0.348 nuScenes检测分数(NDS);在推理时借用监督式框几何将相同检测提升至0.413 mAP / 0.555 NDS,这表明该流程的最大缺陷在于测量精度而非二维检测,而类别混淆和置信度校准在该替换下依然存在。反向来看,基于相同掩码构建的三状态相机见证规则将仅使用LiDAR的监督检测器从0.596提升至0.630 mAP,约为全监督相机融合增益的一半,且无需训练。覆盖分析显示,SAM3能通过正确命名的掩码找到84%的范围内物体;官方指标中失败的类别是命名错误或几何上不宽容的,而非不可见。
英文摘要:
Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.