发表机构
Bosch Center for AI; University of Freiburg; University of Luebeck(博世人工智能中心; 弗莱堡大学; 吕贝克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GhostPoint是一种自监督学习框架,通过实例体素膨胀幻觉生成激光雷达被遮挡区域特征,在nuScenes和Waymo数据集上提升了下游3D检测性能,尤其适用于稀疏扫描和有限标签场景。
AI 中文摘要
从激光雷达点云进行3D目标检测是自动驾驶的核心问题。近期自监督学习(SSL)的进展实现了可扩展的预训练,且能很好地迁移到逐点任务,如语义和全景分割,但迁移到3D检测的效果仍较弱。我们分析了近期的SSL方法,发现大多数目标仅定义在可见表面的激光雷达返回数据上,未对被遮挡和未观测区域进行约束。这种可见表面偏差对逐点预测可能足够,但3D检测需要对缺失结构具有鲁棒性。为解决这一差距,我们提出GhostPoint,这是一种SSL框架,通过新颖的实例体素膨胀,在已发现实例周围的局部邻域中幻觉生成潜在特征。在GhostPoint中,编码器处理观测到的返回数据,额外的预测器从观测上下文推断邻域表示。除了标准的编码器级监督外,我们还对生成邻域中采样的体素引入预测器级监督方案:观测到的(可见/掩码)体素匹配教师编码器目标,而未观测体素匹配教师预测器的幻觉。该设计鼓励学习到的表示显式建模观测返回之外的结构。在nuScenes和Waymo上的广泛评估表明,我们的方法达到了最先进的性能,一致提升了下游3D检测,尤其是在稀疏扫描和有限标签的情况下。
英文摘要
3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer to 3D detection remains weaker. We analyze recent SSL methods and find that most objectives are defined only on measured LiDAR returns from visible surfaces, leaving occluded and unobserved regions unconstrained. This visible-surface bias can be sufficient for point-wise prediction, but 3D detection requires robustness to missing structure. To address this gap, we propose GhostPoint, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation. In GhostPoint, an encoder processes observed returns, and an additional predictor infers neighborhood representations from observed context. In addition to standard encoder-level supervision, we introduce a predictor-level supervision scheme on sampled voxels from generated neighborhoods. Specifically, observed (visible/masked) voxels match teacher-encoder targets, while unobserved voxels match teacher-predictor hallucinations. This design encourages the learned representation to explicitly model structure beyond observed returns. Extensive evaluations on nuScenes and Waymo demonstrate that our method achieves state-of-the-art performance, consistently improving downstream 3D detection, especially under sparse scans and limited labels.
CommentsAccepted by ECCV2026