arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRUISE:面向鲁棒自动驾驶的视觉语言模型引导的不确定性感知跨模态传感器融合

CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque

arXiv 2608.09202首次发表:更新:

发表机构

University of California, Irvine(加利福尼亚大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有不确定性感知融合方法泛化性差的问题,提出CRUISE框架,结合VLM引导的UQ模块与动态自适应机制,实现鲁棒的自动驾驶跨模态传感器融合。

AI 中文摘要

现代自动驾驶汽车配备了相机、激光雷达(LiDAR)和雷达等多种传感器,用于实现全面的环境感知。然而,跨模态特征的鲁棒融合仍是一项关键挑战,因为不同传感器的可靠性在各类真实驾驶场景(包括能见度差和恶劣天气)中存在显著差异。虽然不确定性量化(UQ)可通过让模型优先选择可靠信号缓解该问题,但现有不确定性感知融合方法通常依赖简单的特征级不确定性估计,因此在复杂的分布外场景中往往无法有效泛化。为解决这一局限,我们提出了CRUISE,一种新型不确定性感知跨模态传感器融合框架。CRUISE集成了视觉语言模型(VLM)引导的UQ模块,可生成细粒度的像素级不确定性估计;通过利用VLM丰富的先验知识和出色的上下文推理能力,我们的方法为融合过程提供了高信息量的引导。此外,我们引入了动态自适应机制,明确建模并捕获跨模态依赖关系,确保框架充分利用多传感器输入固有的互补特性。

英文摘要

Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to prioritize reliable signals, existing uncertainty-aware fusion methods typically rely on simple feature-level uncertainty estimates and thus often fail to generalize effectively in complex, out-of-distribution scenarios. To address this limitation, we propose CRUISE, a novel uncertainty-aware cross-modal sensor fusion framework. CRUISE integrates a vision-language model (VLM)-guided UQ module that generates fine-grained, pixel-level uncertainty estimates. By leveraging the VLM's rich prior knowledge and superior contextual reasoning, our approach provides a highly informative guide for the fusion process. Furthermore, we introduce a dynamic adaptive mechanism that explicitly models and captures cross-modal dependencies, ensuring the framework fully exploits the inherent complementary nature of multi-sensor inputs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑