TFTrack:一种用于高效三维点云跟踪的无模板框架
TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking
浏览论文内容
中文总结 AI 辅助
TFTrack提出首个无模板三维点云跟踪框架,仅利用先前边界框中心和大小直接跟踪当前帧,简化运动建模,在KITTI和nuScenes上性能与模板法相当,计算量减半,速度达120 FPS。
中文摘要 AI 辅助
基于激光雷达的三维单目标跟踪(3D SOT)对于机器人感知和导航至关重要,其目标是在稀疏点云中跨帧定位动态物体。现有方法源于二维视觉中的孪生跟踪范式,依赖成本高昂的双输入设计和由模板先验引导的过度运动建模,这阻碍了其效率。我们的深入分析揭示:(i)模板范式是冗余的,因为先前的边界框中心编码了足够的历史上下文;(ii)复杂的运动建模是不必要的,因为几何对齐提供了足够的运动先验。基于上述发现,我们提出了首个无模板跟踪框架(TFTrack)。该新颖框架消除了模板-搜索配对的需求,仅由先前的边界框中心和大小引导,直接对当前帧进行操作。我们将这一范式实例化为三种变体:TFTrack-Voxel、TFTrack-Pillar和TFTrack-Point,以在统一框架下探索不同的三维表示,确保在稀疏和密集场景中的灵活性。在KITTI和nuScenes基准上的大量实验表明,TFTrack与领先的基于模板的跟踪器相比具有竞争力,同时将FLOPs减少约50%,并以约120 FPS的速度运行。通过简化过度复杂的以运动为中心的设计,TFTrack为高效三维点云跟踪建立了一种新的极简范式,为在嵌入式机器人系统(如自动驾驶车辆)中的实时和资源高效部署铺平了道路。代码可在该https URL获取。
英文摘要
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames in sparse point clouds. Existing methods, rooted in the Siamese tracking paradigm from 2D vision, rely on costly dual-input designs and excessive motion modeling guided by template priors, hindering their efficiency. Our in-depth analysis reveals: (i) the template paradigm is redundant, as the previous bounding box center encodes sufficient historical context; (ii) complex motion modeling is unnecessary, as geometric alignment provides adequate motion priors. Based on the above findings, we propose the first Template-Free Tracking framework (TFTrack). The novel framework eliminates the need for template-search pairings and operates directly on the current frame guided solely by the prior bounding box center and size. We instantiate this paradigm into three variants: TFTrack-Voxel, TFTrack-Pillar, and TFTrack-Point, to explore different 3D representations under a unified framework, ensuring flexibility across sparse and dense scenes. Extensive experiments on KITTI and nuScenes benchmarks show that TFTrack is competitive with leading template-based trackers, while reducing FLOPs by approximately 50% and running at approximately 120 FPS. By simplifying overcomplicated motion-centric designs, TFTrack establishes a new minimalist paradigm for efficient 3D point cloud tracking, paving the way for real-time and resource-efficient deployment in embedded robotic systems, such as autonomous vehicles. The code is available at https://github.com/tftrack-anonymous/TFTrack/tree/main.
发表机构
- Stony Brook University(石溪大学)
- Carnegie Mellon University(卡内基梅隆大学)
- Hangzhou Dianzi University(杭州电子科技大学)
- Southeast University(东南大学)
- University of California, Riverside(加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。