发表机构
KAIST; Chung-Ang University(韩国科学技术院; 中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对4D激光雷达标注成本高、难扩展的问题,提出LiDAR-SAM2框架,借助SAM2自动生成标注,大幅降低3D/4D场景理解的标注负担。
AI 中文摘要
4D激光雷达分割的研究进展受限于数据瓶颈:在稀疏点云序列上分配时间一致的标注成本高昂且难以扩展,每个新任务或领域往往都需要全新的密集标注。这催生了一个核心问题:能否在无需任何人工标注的情况下自动生成高质量的激光雷达训练数据?为此,我们提出LiDAR-SAM2框架,将2D视频基础模型SAM2转化为4D激光雷达领域可扩展的监督源。在数据层面,它通过多视图投影和时空聚合,从SAM2的视频掩码自动生成时间一致的激光雷达级标注;在模型层面,定制的模态接口与两阶段学习目标将SAM2的视频分割内核适配于时空激光雷达结构,实现每个物体仅需一次点击即可生成跨序列的一致掩码轨迹。在无人工激光雷达标注的情况下训练的LiDAR-SAM2,在SemanticKITTI数据集上生成的语义和全景标注质量,仅需少量标注点即可接近全人工标注水平,基于这些标注训练的模型性能也接近全真实标注监督的效果,这使LiDAR-SAM2成为可扩展的标注工具,大幅降低了3D和4D场景理解的标注负担。
英文摘要
Progress in 4D LiDAR segmentation is bottlenecked by data. Assigning temporally consistent labels across sparse point cloud sequences is costly and hard to scale, and every new task or domain tends to demand fresh dense annotation. This motivates a simple question of whether high-quality LiDAR training data can be produced automatically, without any human labeling. To this end, we introduce LiDAR-SAM2, a framework that turns a 2D video foundation model, SAM2, into a scalable source of supervision for the 4D LiDAR domain. On the data side, it automatically generates temporally coherent LiDAR-level labels from SAM2 video masks through multi-view projection and spatio-temporal aggregation. On the modeling side, a tailored modality interface and a two-stage learning objective adapt SAM2's video segmentation kernel to spatio-temporal LiDAR structure, so that a single click per object yields a consistent mask track across the sequence. Trained with no human LiDAR annotation, LiDAR-SAM2 produces semantic and panoptic labels on SemanticKITTI that approach the quality of full human annotation from only a few points, and models trained on these labels approach the performance of full ground-truth supervision. This positions LiDAR-SAM2 as a scalable labeling tool that substantially reduces the annotation burden for 3D and 4D scene understanding.
CommentsECCV 2026 Workshop