arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27633cs.CVcs.AIcs.LG

基于边缘端YOLO与RT-DETR的感知路面坑洞深度检测

Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

Md Monjurul Ahsan Prodhan, Md Nour Hossain

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有坑洞检测无法测深的问题,提出感知深度的坑洞检测框架,对比五种模型,在PothRGBD数据集上验证,发现边界框模型存在结构性深度高估偏差,YOLOv8nSeg检测精度最优、YOLOv8n推理最快、RTDETRX置信度最高。

中文摘要 AI 辅助

路面坑洞检测及其严重程度测量仍是城市基础设施管理中的重要挑战,维护不及时会直接导致车辆损坏、道路事故及维修成本攀升。现有自动化方法依赖2D RGB图像,无法测量坑洞的实际深度。本文提出一种感知深度的路面坑洞检测框架,对比五种架构:YOLOv8n、YOLOv8nSeg、YOLOv9t、RTDETRL和RTDETRX,用于基于RGB-D传感器融合的检测与自动深度测量。采用自定义离线增强管线模拟恶劣道路监测场景,所有模型在PothRGBD数据集上训练,按80%训练集、20%验证集划分,以精确率、召回率、mAP@50和mAP@50_95评估。测量深度数据前,用RANSAC地面平面正射校正修正所有深度图的相机倾斜,计算统计量前将传感器零值像素转为NaN。YOLOv8nSeg的mAP@50达0.9556、mAP@50_95达0.6758,结合像素级Dseg算法得到最准确的2.96 cm深度估计;YOLOv8n推理速度最快,为3.6 ms;RTDETRX检测置信度最高,达92.70%。重要发现:即使经完整RANSAC正射校正,与像素级分割掩码相比,边界框模型对坑洞深度的高估为0.16至0.21 cm,证实路面包含偏差是结构性问题而非校准伪影。

英文摘要

Pothole detection and its severity measurement is still an important challenges in urban infrastructure management, where late maintenance directly contributes to vehicle damage, road accidents, and escalating repair costs. Existing automated approaches depend on 2D RGB images and cannot measure physical depth of potholes. In this paper, we present a depthaware pothole detection framework and then compare five architectures: YOLOv8n, YOLOv8nSeg, YOLOv9t, RTDETRL, and RTDETRX for RGB-D sensor fusion-based detection and automated depth measurement. A custom offline augmentation pipeline is used here to simulate adverse road monitoring conditions. All models are trained on the PothRGBD dataset with an 80% training and 20% validation split and evaluated using Precision, Recall, mAP@50, and mAP@50_95. Before measuring the depth data, all depth maps are corrected for camera tilt using RANSAC ground-plane orthorectification and all zero-valued sensor pixels are cast to NaN before any statistic is computed. YOLOv8nSeg achieves the highest mAP@50 of 0.9556 and mAP@50_95 of 0.6758 with the most accurate depth estimate of 2.96 cm with the pixel-precise Dseg algorithm. YOLOv8n achieves the fastest inference at 3.6ms. RTDETRX achieves the highest detection confidence at 92.70%. An important finding is that even after full RANSAC orthorectification, bounding box models overestimate pothole depth by 0.16 to 0.21 cm compared to pixel precise segmentation masks. This confirms that the pavement inclusion bias is structural rather than a calibration artifact.

补充信息

↑