发表机构
Technical University of Munich; Mercedes Benz AG; University of Massachusetts(慕尼黑工业大学; 梅赛德斯-奔驰股份公司; 马萨诸塞大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估了五个视觉语言模型在自动驾驶场景下的鲁棒性,发现视觉损坏对准确性和置信度的影响因模型、数据集和输入设置而异,并验证了推理时方法VEA在部分条件下能提升可靠性。
AI 中文摘要
视觉语言模型(VLM)在自动驾驶领域中的应用日益广泛,涉及场景理解、驾驶推理、决策制定以及端到端驾驶等任务。随着其作用日益突出,确保其鲁棒性和可靠性变得愈发重要。在真实世界条件下,视觉输入可能因传感器缺陷和环境条件而退化,这可能会影响模型的预测及其相关置信度。这种退化在自动驾驶中尤为令人担忧,因为安全关键决策要求模型做出准确预测,并识别其预测可能不可靠的情况。在本工作中,我们评估了五个视觉语言模型(Qwen3.5-9B、Gemma4-E4B、LLaVA-OneVision-7B、DriveFusion/DriveFusionQA-4B 和 NVIDIA Alpamayo-1.5-10B),在四个与驾驶相关的问答数据集上,采用不同的视觉输入设置,包括单帧、多视角、多帧和单目输入。我们的结果表明,视觉损坏的影响因模型、数据集和输入设置而异,准确性和置信度可靠性的变化在不同条件下也有所不同。随后,我们应用视觉证据增强(VEA),这是一种最近的推理时方法,以检验其是否能在退化视觉条件下提高模型可靠性。我们发现,VEA 对某些模型和数据集提升了性能,尽管增益在所有设置中并不一致。
英文摘要
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.