面向中国乡村场景自动驾驶的目标检测:真实-合成数据混合与模型评估的实验研究
Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation
浏览论文内容
中文总结 AI 辅助
针对中国乡村自动驾驶目标检测的数据稀缺与泛化问题,构建真实-合成混合数据集并评估13种主流检测器,发现1:0.5合成数据占比时YOLO11m性能最优,为乡村自动驾驶提供实证支持。
中文摘要 AI 辅助
当前,自动驾驶目标检测模型在复杂的中国乡村交通场景中面临严重的数据稀缺性和泛化挑战。为解决这些局限,我们提出了一种专门针对中国乡村道路的新型真实-合成混合目标检测数据集,并在不同真实-合成数据比例下系统评估了13种主流检测器的性能,从而为乡村自动驾驶场景中的模型选择和数据策略设计提供实证依据。该数据集结合了在河南省尉氏县采集的真实图像,以及通过Unreal Engine生成的参数化虚拟场景。为准确反映乡村交通的独特现实,我们定义了涵盖14类物体的综合体系,包含电动三轮车、低速车辆(LSVs)、路边摊位等区域特有元素。在统一训练协议下,我们评估了YOLOv5、YOLOv8、YOLO11、YOLO26系列及RT-DETR-L共13种主流检测器,设置了全真实基线、1:0.5真实-虚拟混合、1:1混合三种数据配置。实验结果表明,适度注入合成数据(1:0.5比例)可有效提升检测性能,其中YOLO11m的mAP@0.5达到0.758;但更高比例的合成数据(1:1)会引入域偏移,抵消数据扩充的益处。尽管多数模型能可靠识别不同本地车辆,但长尾非标准物体(如摊位、栏杆)仍存在显著感知瓶颈。本研究为模型选择和合成数据策略提供了关键实证依据和新见解,助力自动驾驶感知系统在乡村地区的实际部署。
英文摘要
Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest mAP@0.5 of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.