发表机构
IIT Jodhpur(焦特布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出面向Jetson Orin Nano的精度保留型ORB-SLAM3 GPU实现,通过算法复现ORB前端、结合TensorRT优化回环闭合,在多数据集上验证其精度与原CPU版本相当,实现边缘设备上的高效视觉-惯性SLAM。
AI 中文摘要
低功耗边缘平台上的视觉-惯性SLAM受限于密集特征提取与回环闭合的计算开销。此前ORB-SLAM的GPU移植版本通过近似ORB特征检测器来换取速度,却改变了特征集,进而影响轨迹估计精度。本文针对NVIDIA Jetson Orin Nano提出了精度保留型ORB-SLAM3的GPU实现,其GPU ORB前端通过算法复现参考CPU检测器,实现了94.7%的精确关键点匹配率和99.9%的描述子位匹配率。该工作还通过原生TensorRT使基于CNN的回环闭合具备边缘部署可行性:视觉前端(特征提取)卸载至GPU,建图与优化后端保留在CPU,实现计算与硬件的适配。通过对比四种配置验证精度:GPU管线与未修改的CPU参考版本,分别在Jetson Orin Nano和台式机上运行。在EuRoC数据集上,四种配置的平均绝对轨迹误差(SE(3))均在0.10cm以内,说明GPU移植与硬件变更均未改变轨迹估计结果;在TUM-VI和KITTI数据集上,GPU与CPU的对比结果可复现,表明加速是精度保留而非近似的。该实现在EuRoC上与已发表的ORB-SLAM3性能相当,在6个TUM-VI室内序列中的5个达到亚厘米级精度,在11个KITTI序列中的9个达到亚1%的相对平移误差。对于回环闭合,通用ONNX-Runtime CUDA/TensorRT执行提供程序无法在嵌入式平台上与本文的CosPlace ResNet-50协同,而原生libnvinfer FP16引擎将每次查询的推理时间缩短至2.2ms,实现了180倍加速,因此基于学习的地点识别可在7W设备上与跟踪任务并行运行。在单目-惯性模式下,系统在11个EuRoC序列上维持平均32FPS的帧率。
英文摘要
Visual-inertial SLAM on low-power edge platforms is constrained by the cost of dense feature extraction and loop closure. Prior GPU ports of ORB-SLAM trade accuracy for speed by approximating the ORB detector, altering the feature set and therefore the estimated trajectory. We present an accuracy-preserving GPU implementation of ORB-SLAM3 for the NVIDIA Jetson Orin Nano, whose GPU ORB front end reproduces the reference CPU detector algorithmically to 94.7% exact keypoint agreement and 99.9% descriptor bit agreement. This work also makes CNN-based loop closure edge-viable through native TensorRT. The visual front end (feature extraction) is offloaded to the GPU while the mapping and optimization back end is kept on the CPU, matching each computation to the hardware it suits. The accuracy is verified by comparing four configurations: the GPU pipeline and the unmodified CPU reference, each run on both the Jetson Orin Nano and a desktop. On EuRoC dataset, all four agree to within 0.10cm in mean absolute trajectory error (SE(3)), so neither the GPU port nor the change of hardware shifts the estimated trajectory. The GPU-versus-CPU comparison is reproducible on TUM-VI and KITTI datasets, so the acceleration is accuracy-preserving rather than approximate. The proposed implementation is competitive with published ORB-SLAM3 on EuRoC, attains sub-centimeter accuracy on five of the six TUM-VI room sequences, and reaches sub-1% relative translation error on nine of eleven KITTI sequences. For loop closure, the generic ONNX-Runtime CUDA/TensorRT execution providers are unusable with our CosPlace ResNet-50 on the embedded platform, whereas a native libnvinfer FP16 engine reduces per-query inference to 2.2ms, a 180x speedup. Learned place recognition therefore runs concurrently with tracking on a 7W device. In monocular-inertial mode the system sustains 32FPS mean over the eleven EuRoC sequences.