AI 中文总结
本文提出TriView-YOLO多视图YOLOv12检测器,通过融合三种配准的GPR视图实现软质高含水量土壤的道路空洞检测,在专属野外测试集上取得0.558±0.028的mAP50,性能优于单视图方案,公开及合成训练图像等优化手段无增益。
AI 中文摘要
从探地雷达(GPR)自动检测地下空洞在软质高含水量地层中最为困难,这类地层中导电的水饱和土壤会衰减信号并削弱空洞反射,而这恰恰是空洞最易形成的条件。本文提出TriView-YOLO,一种适用于这类地层道路空洞筛查的多视图YOLOv12检测器。三个配准视图(纵向B扫描、水平C扫描和横截面B扫描)构成9通道输入,由替换YOLOv12主干的TripleInputConv层进行融合;网络其余部分保持不变,仅要求纵向视图上的边界框标注。训练使用1600份经专家验证的野外样本,主要来自泰国曼谷的城市道路勘测,采用车载多通道三维GPR移动测绘系统采集,仅将日本较坚实路基的勘测数据加入训练和验证集。测试集完全来自曼谷的勘测数据,涉及含水量80-140%、地下水位1-2米的软质海相黏土,此前尚无针对该地层的深度学习空洞检测评估报告。在该未增强、仅野外的测试集(按勘测内部随机划分)上,所提模型在3个随机种子下的mAP50为0.558±0.028,浮点运算量为23.6 GFLOPs,单张图像推理耗时3.1毫秒。消融实验显示,移除辅助视图会降低mAP50和召回率,而公开及合成训练图像、DINOv3特征、更大模型规模、COCO预训练均未带来性能提升。
英文摘要
Automated detection of subsurface cavities from Ground Penetrating Radar (GPR) is most difficult in soft, high-water-content ground, where conductive, water-saturated soil attenuates the signal and degrades cavity reflections, yet this is also the condition under which cavities most readily form. This paper proposes TriView-YOLO, a multi-view YOLOv12 detector for road cavity screening in such ground. Three co-registered views (longitudinal B-scan, horizontal C-scan, and cross-section B-scan) form a 9-channel input fused by a TripleInputConv layer that replaces the YOLOv12 stem; the rest of the network is unchanged, and bounding boxes are required on the longitudinal view only. Training used 1,600 expert-verified field samples, principally metropolitan road surveys of Bangkok, Thailand, acquired with a vehicle-mounted multichannel three-dimensional GPR mobile mapping system, with surveys over the firmer subgrades of Japan added to training and validation only. The test set comes exclusively from the Bangkok surveys, over soft marine clay with 80-140% water content and a water table at 1-2 m depth, a ground condition for which no dedicated deep learning cavity-detection evaluation has been reported. On this unaugmented, field-only test set, split randomly within surveys, the proposed model attains mAP50 of 0.558 +/- 0.028 over three seeds at 23.6 GFLOPs and 3.1 ms per image. Ablations show that removing the auxiliary views lowers mAP50 and recall, whereas public and synthetic training images, DINOv3 features, larger model scale, and COCO pretraining bring no gain.