arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05840cs.CVeess.IV

利用监控摄像机图像对道路平面上的道路交通对象进行精确定位

Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery

Jan Gawroński, Witold Czajewski

首次发表
浏览论文内容

中文总结 AI 辅助

针对单目路边监控车辆定位误差大的问题,提出两阶段几何感知定位流水线,经CARLA合成数据训练、DAIR-V2X真实数据微调后,定位精度大幅提升,尤其对远距离车辆效果显著。

中文摘要 AI 辅助

从单目路边监控摄像机进行精确的车辆定位对智能交通系统、交通监控和交通冲突分析具有重要意义。标准方法通常从检测器边界框的中心估计车辆位置,但由于透视畸变和视差,尤其是对于高架摄像机和大型车辆,这会产生较大误差。本文提出了一种两阶段的几何感知定位流水线,用于估计车辆 footprint 在道路平面上的投影。首先,使用基于YOLO26的检测器检测车辆;其次,专用的ResNet34回归网络预测对应于投影车辆底部的四个角点,最终位置计算为预测四边形的几何中心。该方法在CARLA生成的合成数据上进行训练,并在来自DAIR-V2X的真实路边图像上进行微调。在合成和真实数据上的实验表明,与简单的边界框中心定位相比,该方法有明显改进:在DAIR-V2X上,平均图像空间定位误差从31.77像素降至15.30像素,提升了51.8%,中位数误差降至4.29像素;中距离车辆的中位数地面平面误差从5.52米降至0.90米,远距离车辆的中位数地面平面误差从8.67米降至1.84米。结果还表明,检测器边界框周围的上下文信息对几何定位很重要,在远距离车辆以及受强透视畸变和视差影响的几何挑战性案例中,改进最为显著。

英文摘要

Accurate vehicle localization from monocular roadside surveillance cameras is important for intelligent transportation systems, traffic monitoring, and traffic conflict analysis. Standard approaches often estimate vehicle position from the center of the detector bounding box, which can produce large errors due to perspective distortion and parallax, especially for elevated cameras and large vehicles. This paper proposes a two-stage geometry-aware localization pipeline that estimates the projection of the vehicle footprint onto the road plane. First, vehicles are detected using a YOLO26-based detector. Second, a dedicated ResNet34 regression network predicts four corner points corresponding to the projected vehicle base. The final position is computed as the geometric center of the predicted quadrilateral. The method was trained on synthetic data generated in CARLA and fine-tuned on real-world roadside imagery from DAIR-V2X. Experiments on synthetic and real data showed clear improvements over naive bounding-box-center localization. On DAIR-V2X, the mean image-space localization error decreased from 31.77 px to 15.30 px, a 51.8% improvement, while the median error decreased to 4.29 px. Median ground-plane error for medium-range vehicles decreased from 5.52 m to 0.90 m, and for far-range vehicles from 8.67 m to 1.84 m. The results also show that contextual information surrounding the detector bounding box is important for geometric localization. The largest gains were observed for distant vehicles and geometrically challenging cases affected by strong perspective distortion and parallax.

补充信息

↑