将视觉语言模型锚定于驾驶语义:面向可解释推理的多数据集谓词框架
Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning
AI总结:
本文提出确定性多数据集谓词框架,从几何、运动学等证据推导驾驶语义,构建谓词知识图谱,在nuPlan和nuScenes上实现高F1分数,并在NuPlanQA任务中提升视觉语言模型推理准确率。
AI中文摘要:
视觉语言模型越来越多地用于驾驶场景理解,但其输出中表达的语义关系往往难以对照底层交通状况进行验证。本文引入了一个确定性的多数据集谓词框架,该框架从可测量的几何、运动学、时间、地图和交通控制证据中推导驾驶场景语义。数据集特定的接口仅用于恢复所需的场景信息,而谓词定义在nuPlan和nuScenes上保持不变,并在一个共同的谓词知识图谱中具体化。针对每个数据集的200个场景,与人工标注的谓词关系进行定量语义验证,在nuPlan上获得0.94的宏F1分数,在nuScenes上获得0.93的宏F1分数,共享谓词的平均跨数据集差异为0.02。该谓词知识图谱进一步使用冻结的LLaVA-OneVision-7B模型在九个NuPlanQA子任务上进行评估。在九个NuPlanQA子任务中,有七个子任务(包括交通灯(53.2%提升至71.5%)、态势评估(76.2%提升至86.1%)和行动推荐(82.9%提升至89.0%))中,谓词锚定在评估的视觉输入条件下达到了最高准确率。天气/光照基本保持不变(89.4%对比88.8%),这与相应谓词的缺失一致,而仅使用谓词知识图谱输入在九个子任务中的八个优于仅使用元数据输入。结果表明,确定性谓词提供了一致且可追溯的语义表示,并且在理想锚定下,可以减少谓词词汇所覆盖的推理任务对视觉的依赖。
英文摘要:
Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from measurable geometric, kinematic, temporal, map, and traffic-control evidence. Dataset-specific interfaces are used only to recover the required scene information, while predicate definitions remain unchanged across nuPlan and nuScenes and are materialised in a common Predicate Knowledge Graph. Quantitative semantic validation against manually annotated predicate relations on 200 scenarios from each dataset yields macro F1 scores of 0.94 on nuPlan and 0.93 on nuScenes, with an average cross-dataset difference of 0.02 across the shared predicates. The Predicate KG is further evaluated using a frozen LLaVA-OneVision-7B model on the nine NuPlanQA subtasks. Predicate grounding achieves the highest accuracy among the evaluated visual-input conditions in seven of nine NuPlanQA subtasks, including Traffic Light (53.2% to 71.5%), Situation Assessment (76.2% to 86.1%), and Action Recommendation (82.9% to 89.0%). Weather/Lighting remains essentially unchanged (89.4% vs. 88.8%), consistent with the absence of corresponding predicates, while Predicate KG only input outperforms metadata-only input in eight of nine subtasks. The results show that deterministic predicates provide a consistent and traceable semantic representation and, under oracle grounding, can reduce visual dependence for reasoning tasks covered by the predicate vocabulary.