arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动驾驶中基于视觉语言模型的局部共形安全监控

Localized Conformal Safety Monitoring with Vision-Language Models for Autonomous Driving

Luís Marques, Rong Fang, Disha Kamale, Dmitry Berenson

arXiv 2610.02765首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动驾驶安全监控,提出分割标签局部共形预测(SLLCP),在冻结视觉语言模型上实现事后校准,将碰撞检测率从4.6%提升至89.6%。

AI 中文摘要

监控规划的驾驶轨迹需要准确估计与行为受自车运动影响的交通参与者的碰撞可能性。现有的经典方法往往受限于其预测模型的质量。视觉语言模型(VLMs)在推理高层行动的后果方面展现出潜力,但其近似预测不适用于自动驾驶等安全关键应用。共形预测(CP)已成为一种数据驱动的框架,用于量化黑盒模型预测的不确定性。我们提出分割标签局部共形预测(SLLCP),这是一种在冻结的VLM之上的事后校准层,将其不可靠的预测转化为概率校准的安全预测集。我们考虑了安全估计能力如何依赖于观察到的驾驶场景,并引入了一种局部化程序,在计算不确定性阈值时对相关的过往经验进行加权。我们在可交换性假设下提供了标签条件有限样本无分布覆盖保证。在来自未见场景的15k条CARLA轨迹上评估,SLLCP使用Qwen骨干正确标记了89.6%的碰撞轨迹,使用Cosmos骨干标记了88.4%,而基础VLM仅分别标记了4.6%和39.1%的碰撞轨迹。这些结果表明,局部、标签条件的校准可以减少未标记的不安全轨迹。

英文摘要

Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the ego motion. Existing classical approaches are often limited by the quality of their forecasting model. Vision-language models (VLMs) have shown promise in reasoning about the consequences of high-level actions, yet their approximate predictions are unsuitable for safety-critical applications such as autonomous driving. Conformal prediction (CP) has emerged as a data-driven framework for quantifying the uncertainty of black-box model predictions. We propose Split Label-Localized Conformal Prediction (SLLCP), a post-hoc calibration layer over frozen VLMs that transforms their unreliable predictions into probabilistically calibrated safety prediction sets. We consider how the ability to estimate safety can depend on the observed driving scene and introduce a localized procedure that upweights relevant past experience when calculating uncertainty thresholds. We provide label-conditional finite-sample distribution-free coverage under exchangeability. Evaluated over 15k CARLA trajectories from unseen scenarios, SLLCP correctly flags 89.6% of collision-causing trajectories with a Qwen backbone and 88.4% with a Cosmos backbone, while the base VLMs only flagged 4.6% and 39.1% of the collision-causing trajectories, respectively. These results indicate that local, label-conditional calibration can reduce missed unsafe trajectories.

Comments5 pages, 2 figures, 1 table. Extended abstract. Disha Kamale and Dmitry Berenson are joint senior authors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑