发表机构
Tsinghua University; Xiaomi EV; Jilin University(清华大学; 小米汽车; 吉林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言驾驶中检索风险规则不适用场景的问题,提出DRKG与SWRL推理生成场景锚定风险证据,在nuReasoning上NPS提升1.30、NC提升2.76。
AI 中文摘要
检索增强生成(RAG)为视觉-语言驾驶系统提供了外部安全知识,然而,检索到的风险规则可能具有相关性,却不适用于当前场景。接收此类知识的视觉-语言模型(VLM)在决定如何行动之前,必须锚定对象、跨时间绑定实体并验证关系,这使得风险结论的支持依据隐式化。我们通过引入驾驶风险知识图谱(DRKG)和语义网规则语言(SWRL)推理阶段,在VLM决策之前解决这一相关性-适用性差距。结构化感知实例化场景事实,当SWRL规则的前件被联合满足时,这些规则推导出事件和有向风险关系。识别出的事件、绑定的风险关系以及激活规则的语义描述构成了紧凑的证据,这些证据条件化VLM和扩散规划器。在nuReasoning上的配对比较中,我们的方法在规划分数(NPS)上比基于相关性检索的基线提高了1.30分,在非责任碰撞分数(NC)上提高了2.76分。这些增益表明,场景适用的风险证据相对于语义检索的风险知识,能改善安全加权规划。
英文摘要
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.