ESAFusion:通过局部几何互补与多尺度自适应交互的LiDAR-4-D雷达融合用于3-D目标检测
ESAFusion: LiDAR--4-D Radar Fusion via Local Geometric Complementation and Multiscale Adaptive Interaction for 3-D Object Detection
- Shanghai University(上海大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ESAFusion提出证据感知与尺度自适应框架,通过局部几何互补和多尺度交互融合LiDAR与4-D雷达,在VoD数据集上实现最高mAP,提升3-D检测精度与鲁棒性。
AI中文摘要:
LiDAR-4-D雷达融合结合了精确的空间几何与雷达提供的运动和反射率线索,为复杂驾驶环境中的3-D目标检测提供了一种有前景的解决方案。然而,稀疏的雷达观测以及两种模态之间空间采样的差异使得可靠的跨模态互补变得复杂。此外,模态和特征尺度在不同空间区域的相对重要性各不相同,使得自适应融合具有挑战性。为应对这些挑战,我们提出了ESAFusion,一个证据感知和尺度自适应的框架,结合了局部几何互补与多尺度自适应交互。具体来说,我们引入了一个证据感知雷达选择(ERS)模块,利用运动和观测质量证据抑制雷达杂波,同时保留前景置信度以供后续融合。然后,柱级互补编码器(PCE)利用相邻LiDAR柱的局部几何支持,在空间采样不匹配的情况下改进跨模态互补。我们进一步设计了一个尺度内和尺度间自适应融合(ISAF)模块,以自适应调整不同模态和特征尺度在鸟瞰图(BEV)空间中的贡献。在View-of-Delft(VoD)数据集上的大量实验表明,ESAFusion在比较方法中实现了最高的平均精度(mAP),在完整标注区域达到74.60%,在驾驶走廊达到88.89%。它还在两个区域中获得了最高的骑行者平均精度(AP),同时以19.23 FPS运行。在VoD-Fog上的评估进一步证明了在LiDAR观测逐渐退化的情况下的鲁棒性。源代码将在此https URL公开提供。
英文摘要:
LiDAR--4-D radar fusion combines accurate spatial geometry with motion and reflectivity cues from radar, offering a promising solution for 3-D object detection in complex driving environments. However, sparse radar observations and differences in spatial sampling between the two modalities complicate reliable cross-modal complementation. Moreover, the relative importance of modalities and feature scales varies across spatial regions, making adaptive fusion challenging. To address these challenges, we propose ESAFusion, an evidence-aware and scale-adaptive framework that combines local geometric complementation with multiscale adaptive interaction. Specifically, we introduce an Evidence-Aware Radar Selection (ERS) module to suppress radar clutter using motion and observation-quality evidence while retaining foreground confidence for subsequent fusion. Then, the Pillar-Level Complementary Encoder (PCE) improves cross-modal complementation under mismatched spatial sampling using local geometric support from neighboring LiDAR pillars. We further design an Intra- and Inter-Scale Adaptive Fusion (ISAF) module to adaptively adjust the contributions of different modalities and feature scales in bird's-eye-view (BEV) space. Extensive experiments on the View-of-Delft (VoD) dataset show that ESAFusion achieves the highest mean average precision (mAP) among the compared methods, reaching 74.60% in the Entire Annotated Area and 88.89% in the Driving Corridor. It also attains the highest average precision (AP) for Cyclist among these methods in both regions while running at 19.23 FPS. Evaluations on VoD-Fog further demonstrate robustness under progressively degraded LiDAR observations. The source code will be made publicly available at https://github.com/SenJieHu549/ESAFusion.