arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16490cs.CV

面向实时且自适应的激光雷达场景补全

Towards Real-Time and Adaptable LiDAR Scene Completion

Azhar Hussian, Martin Vossiek, Vasileios Belagiannis

首次发表
浏览论文内容

中文总结 AI 辅助

提出RapidLiDAR方法,以自适应初始化模块和多尺度重建模块实现激光雷达场景补全,速度达0.1秒/帧,比现有最快方法快2.3倍,性能与SOTA相当。

中文摘要 AI 辅助

激光雷达场景补全是自动驾驶中3D感知的关键组成部分,该场景必须实时补全才能用于下游任务。现有方法通常遵循“初始化-优化”范式,先构建场景的粗略初始化结果,再将其优化为完整的3D几何结构。生成模型速度较慢,因为它们会迭代地将随机高斯噪声优化为场景;而非生成方法用固定噪声尺度对部分场景进行扰动,这限制了对大间隙和遮挡区域的覆盖,且需要针对每种新的传感器配置手动重新校准。我们提出RapidLiDAR,一种激光雷达场景补全方法,该方法将初始化本身视为一个学习的、数据驱动的组件。我们提出自适应初始化模块,该模块为每个部分输入点预测空间变化的位移,将部分观测扩展为适配局部几何的粗略场景初始化,无需手动调整噪声。为将该粗略初始化优化为完整且连贯的场景,我们还提出多尺度重建模块,该模块通过查询从输入扫描构建的多尺度3D体素和2D BEV特征图,进一步优化点位置。通过将最远点采样和k近邻搜索等点邻域算子替换为基于体素和BEV的特征提取,我们的架构速度更快,且可按设计处理不同输入分辨率。在SemanticKITTI和KITTI-360上的实验表明,我们的方法达到了与现有最先进方法相当的补全性能,同时在0.1秒内完成完整场景补全,比现有最快方法快2.3倍,这匹配了典型汽车激光雷达传感器10Hz的采集速率,向实时激光雷达场景补全迈出了一步。

英文摘要

LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse initialization of the scene is first constructed, then refined into complete 3D geometry. Generative models are slower because they iteratively refine random Gaussian noise into the scene, while non-generative methods perturb the partial scene with a fixed noise scale, which limits coverage of large gaps and occluded regions and requires manual recalibration for each new sensor configuration. We present RapidLiDAR, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component. We propose an adaptive initialization module that predicts a spatially varying displacement for each partial input point, expanding the partial observations into a coarse scene initialization adapted to the local geometry, without requiring manual noise tuning. To refine this coarse initialization into a complete and coherent scene, we additionally propose a multi-scale reconstruction module that further refines point positions by querying multi-scale 3D voxel and 2D BEV feature maps constructed from the input scan. By replacing point-neighborhood operators such as farthest point sampling and $k$-nearest neighbor search with voxel- and BEV-based feature extraction, our architecture is faster and can handle different input resolutions by design. Experiments on SemanticKITTI and KITTI-360 show that our method achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method. This matches the 10 Hz acquisition rate of typical automotive LiDAR sensors, taking a step toward real-time LiDAR scene completion.

发表机构

  • Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑