发表机构
Infraplus Co., Ltd.(因弗拉普拉斯有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对路面病害检测的小目标前提提出质疑,构建YOLO26-RD检测器,通过优化检测层级等设计,在7618张图像数据集上提升了mAP50,为路面病害检测提供了测量驱动的优化思路。
AI 中文摘要
自动化路面病害检测通常被视为小目标问题,因此催生了高分辨率P2/4检测头与无损下采样方案。我们提出YOLO26-RD,一种基于YOLO26构建的端到端(非极大值抑制NMS-free)检测器,包含两个轻量新型模块:LearnableContrast,一种参数为494的可微分CLAHE类似物,可在网络内每个图块中自适应调整对比度;EdgeSPD,一种Sobel门控空间到深度下采样器,相比SPD-Conv仅增加2个参数。我们在含7618张图像的道路调查数据集(涵盖龟裂、线状裂缝、修补区域)上对该设计开展数据优先审计。审计结果否定了小目标前提:92%的实例属于COCO大型目标,且线状裂缝是极端长宽比结构(中位数10:1),其难点在于灵敏度而非定位。基于此分析,我们移除P2检测层级但在融合路径中保留P2特征,这使mAP50较完整YOLO26-RD模型提升2.8个百分点,同时将每轮训练时间缩短8%。在640×640分辨率下从头训练,我们的最佳筛选配置在验证集上达到0.787的mAP50,而项目基线(未调整配方的标准YOLO26-s)为0.771;最稀有类别(修补区域)的每类增益最大,较基线提升2.6个百分点,较相同配方下未修改的YOLO26-RD对照组提升9.4个百分点。故障模式分解进一步将瓶颈类别(裂缝,所有测试架构中约为0.74)的残留误差归因于裂缝方向与长度的训练/验证分布偏移、640×640分辨率下的亚像素裂缝宽度及标签不完整性,这些因素是任何架构变更都无法解决的。我们认为,对于路面图像,测量驱动的减法优于模块累加,且我们会发布审计协议与模型。
英文摘要
Pavement distress detectors are conventionally specialised for small objects, typically by adding a stride-4 detection head and replacing strided convolution with space-to-depth downsampling. This paper tests that premise against the annotation geometry of region level survey imagery and finds it fails: 1.28% of instances are small at 640 resolution while 70.37% are large, yet a stride-4 level would claim 75.3% of anchors, and complete misses rather than localisation errors dominate baseline failures. YOLO26-RD therefore reallocates the anchor budget, retaining the stride-4 branch as neck features but carrying no detection level there, and adds LearnableContrast, a 494 parameter per tile correction learned from the detection loss and active at inference, and EdgeSPD, a lossless space-to-depth downsampler gated by a fixed Sobel prior. Fifteen models were trained from scratch under one recipe, five scales each of YOLO26-RD and of matched YOLO26 and YOLOv12 families. Averaged over scales YOLO26-RD returns 0.790 mAP50 and 0.482 mAP50-95 against 0.776 and 0.471 for YOLO26 and 0.755 and 0.468 for YOLOv12; it exceeds both on mAP50 at every scale from s upward, and at m, l and x it leads on both metrics, twelve pairwise comparisons decided without exception. YOLO26-RD-l is the best of the fifteen at 0.809 mAP50 and 0.497 mAP50-95, improving on the YOLO26 reference by 0.031 and 0.030 and leading all six per class entries; every arm of a module ablation also exceeds that reference. The margin is thus a property of the architecture rather than of one tuned configuration, though three of the twelve margins lie inside the dataset 0.015 resolution limit and the held out split reproduces the ordering against YOLO26 but not YOLOv12 at scale x. As a TensorRT FP16 engine the released model sustains 98 frames per second on an entry level accelerator, against the 21 needed at 100 km/h.