arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CalibBEV:基于鸟瞰图对齐的激光雷达-相机标定方法

CalibBEV: LiDAR-Camera Calibration via BEV Alignment

Filippo D'Addeo, Lorenzo Cipelli, Adriano Cardace, Emanuele Ghelfi, Andrea Zinelli, Massimo Bertozzi

arXiv 2608.02309首次发表:更新:

发表机构

University of Bologna; University of Parma; Stanford University; VisLab srl; Ambarella Inc.(博洛尼亚大学; 帕尔马大学; 斯坦福大学; VisLab有限公司; 安霸公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CalibBEV是一种激光雷达-相机标定方法,通过两步BEV对齐结合CLIP对比损失实现跨模态统一,在KITTI、nuScenes基准上大幅降低了标定误差,达到最优性能。

AI 中文摘要

我们提出了CalibBEV,一种用于激光雷达-相机标定的新型鸟瞰图(Bird's Eye View, BEV)对齐方法。该方法将激光雷达与相机数据统一到共享的3D空间表示中,实现了精准且鲁棒的跨模态标定。CalibBEV利用各模态的专用架构提取传感器专属的BEV特征,并通过两步对齐过程估计标定矩阵。第一步,我们通过直接从BEV特征回归得到粗略标定矩阵来执行隐式对齐;为简化该对齐过程,我们借鉴CLIP的对比损失,强制各模态BEV表示之间的语义一致性,引导两个网络向统一特征空间收敛。第二步,我们利用BEV公式显式对齐两个模态的特征,将初始粗略估计细化为最终更精准的标定矩阵。CalibBEV的性能显著优于现有点-像素匹配方法,达到了当前最优的标定精度;在KITTI和nuScenes基准上,相比先前方法,该方法将相对旋转误差(Relative Rotation Error, RRE)分别降低51%和68%,相对平移误差(Relative Translation Error, RTE)分别降低80%和91%。

英文摘要

We present CalibBEV, a novel Bird's Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross-modal calibration. CalibBEV extracts sensor-wise BEV features from each modality using domain-specific architectures and estimates the calibration matrix through a two-step alignment process. First, we perform an implicit alignment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce semantic consistency between BEV representations across modalities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the initial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to-pixel matching methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.

Journal refProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2026. p. 4345-4354

DOI:10.1109/WACV61042.2026.00423

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑