Vernata:激光雷达点表示的自监督学习
Vernata: Self-Supervised Learning of LiDAR Point Representations
- Robotics and AI Institute(机器人与人工智能研究院)
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出基于Sonata架构的Vernata自监督学习框架,通过三项扩展优化激光雷达点云表示,在多数据集上相比基线取得显著性能提升,模态减少场景下仍保持竞争力
AI中文摘要:
激光雷达是户外作业机器人的主要感知模态,但该领域深度学习模型的性能受限于标注数据的稀缺性,这是3D标注成本高昂导致的直接结果。自监督学习可通过从未标注数据中学习通用特征来缓解数据稀缺问题。本研究提出一种面向户外激光雷达点云的自监督多模态多教师蒸馏框架,该框架基于Sonata架构构建,引入了Vernata,包含三项扩展:稀疏视图增强以提升对不同点密度的鲁棒性、内存库机制以稳定资源受限的训练过程、利用密集高分辨率2D图像特征的跨模态蒸馏以实现细粒度语义引导。我们在GrandTour、TartanGround、Waymo数据集及自研机器人平台采集的数据上评估了所提方法,实验结果表明其相比Sonata基线有显著性能提升:在TartanGround上的mIoU得分为54.7(提升5.9个百分点,增幅12.1%),在Waymo上的mIoU得分为57.1(提升7.3个百分点,增幅14.7%)。最后,我们验证了该自监督方法在模态减少场景(缺失颜色或法向量)下仍保持较强性能,在对应数据集上分别取得49.4和50.2的竞争力mIoU得分。
英文摘要:
LiDAR serves as a primary sensing modality for robots operating in outdoor environments. However, the performance of deep learning models in this domain is severely limited by the scarcity of labeled data, a direct result of the high cost of 3D annotation. Self-supervised learning addresses this scarcity by learning general-purpose features from unlabeled data. In this work, we present a multi-modal, multi-teacher distillation framework for self-supervised learning on outdoor LiDAR point clouds. Building upon the Sonata architecture, we introduce Vernata, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guidance. We evaluate our method on the GrandTour, TartanGround, and Waymo datasets, as well as data collected from our own robotic platforms. Our experiments demonstrate a significant performance improvement over Sonata baselines, yielding mIoU scores of 54.7 on TartanGround (+5.9 points, +12.1%) and 57.1 on Waymo (+7.3 points, +14.7%). Finally, we show that the self-supervised approach maintains strong performance even in reduced-modality settings (lacking color or normals), achieving competitive mIoU scores of 49.4 and 50.2 on the respective datasets.