arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29836cs.CVcs.RO

SplatLabel:通过4D高斯泼溅进行伪标注

SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting

  • Perciv AI
  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Nitya Nanvani, Andras Palffy, Holger Caesar

AI总结:

SplatLabel利用4D高斯泼溅自动生成LiDAR分割与占用伪标签,通过时间流形跟踪动态物体,无需3D框标注,并在SemanticKITTI上超越现有方法。

AI中文摘要:

虽然2D视觉基础模型为自动化3D语义伪标注提供了一条途径,但将这些先验知识转化为稳健的3D表示通常需要复杂的启发式方法或多模型集成。我们提出了SplatLabel,一种自动化流水线,利用4D高斯表示来提取带有预测置信度的LiDAR分割结果,以及任意体素分辨率下的语义占用网格。其核心在于,SplatLabel通过一个显式的时间流形来处理动态环境,该流形对单个3D图元的轨迹和生命周期进行建模。这使得系统能够准确跟踪移动物体,并严格定义物体出现和消失的时间,完全消除了对预标注3D边界框的需求。为了稳健地支持这种动态跟踪,该表示以结构和语义先验为基础:我们通过虚拟深度图集成360度LiDAR来引导未观测区域中的场景几何,并且不依赖特定领域的提示工程,而是直接从2D模型蒸馏连续的软概率,以从本质上解决时间和空间上的语义模糊性。最后,为了准确反映精确率与召回率之间的现实权衡,我们将伪标注评估重新定义为一种选择性分类任务,使用广义的风险-召回率度量。在SemanticKITTI上的实验表明,SplatLabel在多个召回率水平上持续优于最先进的基线,为3D LiDAR分割和占用预测建立了一个高度稳健的框架。

英文摘要:

While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract LiDAR segmentation with predictive confidence, as well as semantic occupancy grids at arbitrary voxel resolutions. At its core, SplatLabel handles dynamic environments through an explicit temporal manifold that models the trajectories and lifespans of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-annotated 3D bounding boxes. To robustly support this dynamic tracking, the representation is grounded by structural and semantic priors: we guide scene geometry in unobserved regions by integrating 360-degree LiDAR via virtual depth maps, and rather than relying on domain-specific prompt engineering, we directly distill continuous soft probabilities from 2D models to inherently resolve semantic ambiguities over time and space. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.

↑