AI 中文总结
针对室内移动机器人在严格延迟和内存限制下的LiDAR感知难题,IndoorBEV通过高度感知BEV表示和几何条件特征融合,以0.6M参数实现实时语义地图与定向框输出,在精度、延迟和内存间取得良好平衡。
AI 中文摘要
高效的室内LiDAR感知具有挑战性,因为移动机器人必须在严格的延迟和内存限制下理解杂乱的三维环境。现有的基于点和基于体素的方法通常会产生大量的计算开销,而传统的鸟瞰图(BEV)表示虽然提高了效率,但以丢弃垂直几何信息为代价。我们提出了IndoorBEV,一种轻量级LiDAR感知框架,通过高度感知的BEV表示和几何条件特征融合来缓解这一权衡。IndoorBEV使用统计高度特征和多频率高度编码来总结每个BEV单元中的垂直点分布,使得信息丰富的三维线索能够通过二维卷积高效处理。然后,一个轻量级编码器将互补的几何特征与多尺度局部表示和紧凑的全局场景上下文集成。解耦的密集预测头联合生成语义BEV地图和定向物体边界框。IndoorBEV仅包含0.6M参数,需要2.3 MB的模型存储。在NVIDIA AGX Orin上,每次推理使用21.52 MB的GPU内存,在200毫秒的感知截止时间下实现平均延迟169.6毫秒,截止时间错过率为1.8%。在模拟场景、真实机器人扫描和开源室内点云数据集上的评估表明,在感知精度、延迟和内存消耗之间取得了有利的权衡。这些结果表明,在紧凑的BEV表示中显式编码垂直几何为资源高效的室内LiDAR感知提供了一种有效方法。
英文摘要
Efficient indoor LiDAR perception is challenging because mobile robots must understand cluttered three-dimensional environments under strict latency and memory constraints. Existing point-based and voxel-based methods often incur substantial computational overhead, whereas conventional bird's-eye-view (BEV) representations improve efficiency at the cost of discarding vertical geometric information. We present IndoorBEV, a lightweight LiDAR perception framework that mitigates this tradeoff through a height-aware BEV representation and geometry-conditioned feature fusion. IndoorBEV summarizes the vertical point distribution in each BEV cell using statistical height features and multi-frequency height encoding, allowing informative three-dimensional cues to be processed efficiently by two-dimensional convolutions. A lightweight encoder then integrates complementary geometric features with multi-scale local representations and compact global scene context. Decoupled dense prediction heads jointly produce semantic BEV maps and oriented object bounding boxes. IndoorBEV contains only 0.6M parameters and requires 2.3 MB of model storage. On an NVIDIA AGX Orin, it uses 21.52 MB of GPU memory per inference and achieves a mean latency of 169.6 ms under a 200 ms perception deadline, with a deadline miss ratio of 1.8\%. Evaluations on simulated scenes, real-world robot scans, and an open-source indoor point-cloud dataset demonstrate a favorable tradeoff among perception accuracy, latency, and memory consumption. These results indicate that explicitly encoding vertical geometry within a compact BEV representation provides an effective approach to resource-efficient indoor LiDAR perception.