arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29194cs.CV

RLG-TPV:用于相机-雷达3D目标检测的雷达与激光雷达引导的三视角视图融合

RLG-TPV: Radar- and LiDAR-Guided Tri-Perspective View Fusion for Camera-Radar 3D Object Detection

  • Ozyegin University(厄兹伊金大学)

机构由 AI 辅助整理,请以论文原文为准。

Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates

AI总结:

提出RLG-TPV多模态TPV框架,融合雷达与激光雷达几何引导,结合射线引导注意力、多普勒时间融合等技术,在nuScenes验证集上较CRN基线显著降低方向和速度误差,提升相机-雷达3D目标检测性能。

AI中文摘要:

三视角视图(Tri-Perspective View,TPV)表示通过顶视图、侧视图和前视图特征平面描述3D场景结构,但现有TPV提升方法主要基于相机,导致采样图像证据在投影相机射线上的深度模糊。我们提出RLG-TPV,一种用于相机-雷达3D目标检测的多模态TPV框架,其中雷达和训练时的激光雷达在表示构建过程中提供互补几何引导。射线引导的可变形注意力提升利用激光雷达监督的相机深度概率和雷达视锥占用率对采样图像特征进行加权,而雷达在提升前进一步优化深度分布。由于常规雷达提供的高程信息有限,训练期间由激光雷达导出的类别占用目标监督侧视图和前视图平面;推理时移除相应的头,因此部署仅需相机和雷达。对于时间聚合,多普勒引导的时间融合利用以测量的雷达径向速度为锚定的运动场对齐过去特征,门控机制限制无支持运动区域的变形。此外,雷达散射截面(RCS)感知的雷达散射使雷达证据能在由雷达散射截面条件决定的空间邻域内传播。在nuScenes验证集上,RLG-TPV达到0.4981的平均精度(mAP)和0.5959的nuScenes检测分数(NDS),与已发表的CRN基线相比,方向误差降低31.9%,速度误差降低30.7%。消融研究表明,射线级几何引导是最终性能的主要贡献因素。

英文摘要:

Tri-Perspective View (TPV) representations describe 3D scene structure through top, side, and front feature planes, but existing TPV lifting is primarily camera-based, leaving the depth of sampled image evidence ambiguous along projected camera rays. We propose RLG-TPV, a multimodal TPV framework for camera-radar 3D object detection in which radar and training-time LiDAR provide complementary geometric guidance during representation construction. A ray-guided deformable-attention lift weights sampled image features using LiDAR-supervised camera depth probabilities and radar frustum occupancy, while radar additionally refines the depth distribution before lifting. Because conventional radar provides limited elevation information, LiDAR-derived class-occupancy targets supervise the side and front planes during training; the corresponding heads are removed at inference, so deployment requires only cameras and radar. For temporal aggregation, Doppler-guided temporal fusion aligns past features using a motion field anchored by measured radar radial velocity, with gating that limits warping in regions without supported motion. An RCS-aware radar scatter further allows radar evidence to spread over spatial neighborhoods conditioned on radar cross section. On the nuScenes validation set, RLG-TPV achieves 0.4981 mAP and 0.5959 NDS, reducing orientation and velocity error by 31.9\% and 30.7\% relative to the published CRN baseline. Ablation studies show that ray-level geometric guidance is a major contributor to the final performance.

↑