arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FogDrive:用于分级雾天感知的多模态合成驾驶数据集

FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog

Vansh Panwar

arXiv 2607.22698首次发表:更新:

发表机构

Indian Institute of Technology, Guwahati(印度理工学院古瓦哈提分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FogDrive是用于分级雾天感知的多模态合成驾驶数据集,基于CARLA模拟器构建,含多种传感器数据,模拟不同密度雾。通过跨两种范式建立基线基准,得出训练中混合多密度雾可收紧3D边界框几何形状等见解,将开源加速多模态研究。

AI 中文摘要

恶劣天气下的感知仍然是可靠自动驾驶的关键瓶颈,现有基准缺乏评估稳健传感器融合所需的系统多模态对齐。真实世界天气数据集存在采集不受控和单级、未校准条件的问题,而合成数据集要么仅针对相机恢复,要么缺乏对“去雾然后检测”管道进行基准测试所需的配对清晰和有雾结构。我们提出了FogDrive,这是一个经过严格校准的多模态自动驾驶数据集,它连接了以数据为中心的工程和强大的机器学习。FogDrive基于CARLA模拟器构建,包含660个场景(约133k个完全注释的帧,白天/黑夜比例为50:50),跨越四个同步摄像头(RGB、深度、语义分割)、一个激光雷达和语义激光雷达对以及前雷达。在三个校准能见度密度(160m、100m、50m)下,在相机通道(科斯米德模型)和激光雷达通道(比尔-朗伯定律)上独立模拟物理上一致的雾。每个场景都有四个匹配的变体(清晰加上三个分级雾级别),带有交叉校准二维和三维边界框。对8k图像进行的基于语义分割的质量审核验证了40m内车辆注释的精度为95.1%,召回率超过99%。我们使用跨两种范式(3D多模态融合和2D图像恢复)的先进架构(TransFusion、BEVFusion、YOLOv8-m)建立了基线基准。这些产生了关键的以数据为中心的见解:在训练期间混合多密度雾可在不增加数据缩放成本的情况下收紧三维边界框几何形状,而在二维管道中,图像质量指标(PSNR、SSIM)证明是下游检测性能的较差预测指标。FogDrive将与我们的数据生成框架一起完全开源,以加速稳健的多模态研究。

英文摘要

Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark "defog-then-detect" pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑