发表机构
School of Computing; Institute of Science Tokyo(计算机学院; 东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlashReg是面向GPU的点云配准算法,通过优化图构建与3-团搜索,在保持相当召回率的同时降低延迟、减少内存占用,适用于车载感知的高吞吐量配准后端。
AI 中文摘要
基于图的点云配准通过识别几何一致的对应集实现高鲁棒性,但构建二阶兼容性图和枚举候选团的过程计算与内存密集度高。本文提出FlashReg,一种面向GPU的对应关系到位姿估计器,避免构建密集的带分二阶图。其快速一阶和二阶图(FFSOG)构建从二进制一阶图直接构建容量受限的稀疏二阶图。经数据流优化的三节点团(3-团)搜索从紧凑的逐行候选池中选择枢轴,并通过排序后的稀疏邻域交集枚举三元组。在室内和室外基准测试中,FlashReg在配准召回率相当的情况下,相比TurboReg将对应关系到位姿的延迟降低2-3倍,且在嵌入式GPU上仅使用其峰值分配张量内存的约50%。这些结果使FlashReg适合作为车载感知流水线中的高吞吐量配准后端。
英文摘要
Graph-based point cloud registration achieves high robustness by identifying geometrically consistent correspondence sets, but constructing second-order compatibility graphs and enumerating candidate cliques remain compute- and memory-intensive. This work presents FlashReg, a GPU-oriented correspondence-to-pose estimator that avoids materializing the dense scored second-order graph. Its Fast First- and Second-Order Graph (FFSOG) construction builds a capacity-bounded sparse second-order graph directly from the binary first-order graph. A dataflow-optimized three-node clique (3-clique) search then selects pivots from compact per-row candidate pools and enumerates triples through sorted sparse-neighborhood intersections. Across indoor and outdoor benchmarks, FlashReg reduces correspondence-to-pose latency by 2--3x relative to TurboReg at comparable registration recall, while using about 50% of its peak allocated tensor memory on an embedded GPU. These results make FlashReg suitable as a high-throughput registration backend within onboard perception pipelines.
Comments12 pages