发表机构
School of Mechanical Engineering, Beijing Institute of Technology; National Engineering Research Center of Electric Vehicles, Beijing Institute of Technology; School of Mechanical Engineering, Southeast University(北京理工大学机械工程学院; 北京理工大学电动车辆国家工程研究中心; 东南大学机械工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自动驾驶高级规划问题,提出认知双过程规划框架,用S-CoT模式表示场景知识,通过自动数据引擎、视觉仲裁器及验证器实现自适应推理和规则验证,提升规划准确性、逻辑一致性并降低延迟,明确性能下降条件。
AI 中文摘要
自动驾驶的高级规划是一项知识密集型工程决策任务,需要准确的场景理解、及时的推理和内部一致的行动选择。视觉语言模型(VLMs)虽能使中间推理明确,但在实际规划器中的应用受成本高的结构化监督、常规场景中不必要推理以及生成的理由与驾驶行动可能不一致的限制。我们提出了一个认知双过程规划框架,用机器可解析的结构化思维链(S-CoT)模式表示与规划相关的场景知识。通过自动数据引擎生成S-CoT监督,用轻量级视觉仲裁器根据场景复杂性进行路由选择,对慢路径输出用基于规则的验证器检查一致性并提供可验证奖励。在手动审核和测试样本中取得了较好结果,还识别出路由和规划性能下降的条件。这些结果表明,明确的场景知识可通过自适应推理和基于规则的验证来支持高级VLM规划决策。
英文摘要
High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely inference, and internally consistent action selection. Vision-language models (VLMs) can make intermediate reasoning explicit, but their use in deployed planners is constrained by costly structured supervision, unnecessary reasoning in routine scenes, and possible inconsistencies between generated rationales and driving actions. We present a cognitive dual-process planning framework that represents planning-relevant scene knowledge in a machine-parsable structured chain-of-thought (S-CoT) schema. An automated data engine integrates perception foundation models, critical-path filtering, and an expert VLM to generate S-CoT supervision without manual annotation of individual rationales. A lightweight visual Arbiter estimates scene complexity from multilevel vision-encoder features before language decoding and routes each input to either fast meta-action prediction or slow structured reasoning. For slow-path outputs, a deterministic rule-based validator checks whether the parsed S-CoT fields are consistent with the final meta-action and provides verifiable rewards for Group Relative Policy Optimization (GRPO). In a 195-scene manual audit, the generated annotations achieve 91.8\% CoT accuracy and a 98.5\% Logical Consistency Score (LCS). On 574 manually verified NAVSIM test samples, the planner achieves 80.14\% planning accuracy and 97.20\% LCS while reducing average latency by 17.39\% relative to applying slow reasoning to every scene. Evaluation on external long-tail subsets further identifies conditions under which routing and planning performance degrade. Together, these results show how explicit scene knowledge can be operationalized through adaptive reasoning and rule-based verification to support high-level VLM planning decisions.