arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VOLA:通过基于VLM的语义属性预测改进开放世界驾驶

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

Yuchen Zhang, Yuan Gao, Sebastian Schmidt, Johannes Betz

arXiv 2608.11777首次发表:更新:

发表机构

Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VOLA模型基于VLM图像token预测驾驶相关属性,在开放世界驾驶场景中,对训练外新障碍物的易损性等级召回率优于纯视觉分割器和提示式VLM分割器。

AI 中文摘要

真实世界驾驶属于开放场景:车辆可能遇到掉落的床垫、鹿或训练数据之外的其他物体。仅命名这些物体是不够的,系统必须知道如何对待每个区域:能否从其上驶过,碰撞的严重程度如何?因此,我们将场景感知从类别标签转向与动作相关的密集属性,其中每个像素根据其对运动的影响方式进行标注,而非物体名称。我们用两个有序属性实例化这一通用框架:7级可行驶性和5级易损性。我们直接读取Qwen3.5图像token的隐藏状态作为空间语义表示,一个轻量的边界感知解码器随后将这种粗糙的token网格转换为高分辨率的完整属性图。整个过程既不需要自回归文本生成,也不需要外部掩码模型如SAM。我们在CARLA中构建的密集属性标签上进行训练,并测试其向真实场景和训练中从未见过的新障碍物的迁移能力。我们与在相同属性上训练的纯视觉分割器以及提示式VLM分割器进行比较。我们的模型在熟悉类别上与强大的纯视觉分割器表现相当,且提升了向真实开放世界异常的迁移能力,达到69.4%的平均易损性等级召回率,而最佳纯视觉基线为57.1%,最佳提示式VLM基线为53.9%。这些结果表明,VLM图像token为将驾驶属性迁移至训练词汇之外的物体提供了有用的语义线索。

英文摘要

Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.

Comments16 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑