发表机构
Department of Computer Science and Engineering, The Chinese University of Hong Kong; Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系; 香港中文大学医学智能与扩展现实研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究恶劣天气下视觉语言模型在多模态输入时的表现,引入ObsDriveBench基准测试,通过可观测性元标注等构建含多类问题的测试集,实验发现现有模型性能降,还引入ObsDrive模型提升其在多方面的鲁棒性。
AI 中文摘要
恶劣天气下的自动驾驶仍是关键挑战,现有视觉语言基准测试主要在标准条件、合成损坏或单模态下评估。因此,尚不清楚视觉语言模型在多模态输入的真实恶劣天气下的表现。关键困难在于环境可观测性下降,多模态观测变得不可靠且跨模态不一致。为此引入ObsDriveBench,一个用于恶劣天气自动驾驶的真实多模态基准测试,设计了三个能力维度。通过可观测性元标注等构建基准测试,含超14k训练和13k测试问题。实验显示现有模型性能下降,还引入ObsDrive模型提升了鲁棒性。数据集和评估代码将发布。
英文摘要
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbf{observability awareness}, \textbf{spatial reliability}, and \textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \href{https://github.com/russellyq/ObsDriveBench}{\texttt{ObsDriveBench}}.