发表机构
Fudan University; Yinwang Intelligent Technology Co., Ltd; Fuzhou University; Sun Yat-Sen University; Huawei Technology(复旦大学; 银望智能科技有限公司; 福州大学; 中山大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CAR-VLA,联合场景复杂度与动态风险,映射为三种推理模式,通过渐进式监督学习和推理增强强化学习训练,在多个基准上实现竞争力驾驶性能。
AI 中文摘要
现有的驾驶视觉-语言-动作(VLA)模型自适应推理方法主要关注是否进行推理,而忽略了推理在不同驾驶情境中应如何差异化。我们的关键洞察是,虽然场景复杂度决定了推理深度,但动态风险对于在时间紧迫的情境中决定如何推理同样至关重要。因此,我们提出CAR-VLA,一个统一的驾驶VLA模型,它联合考虑场景复杂度和动态风险,以指导推理的深度、紧迫性和焦点。CAR-VLA将四种复杂度-风险类别映射到三种推理模式:\textit{快速直觉}(Fast Intuition)用于简单低风险场景中的直接轨迹生成,\textit{慢思考}(Slow Thinking)用于复杂低风险场景中的深思熟虑推理,以及\textit{反射响应}(Reflex Response)用于高风险场景中的紧凑、以危险为中心的推理,无论复杂度如何。反射响应并非仅仅缩短思考时间,而是将推理集中在最关键的危害和即时的安全响应上。我们通过渐进式监督学习训练CAR-VLA,将场景评估、推理模式选择和轨迹生成联系起来,随后进行推理增强的强化学习以提高驾驶质量和推理行为。在NAVSIM v1(91.1 PDMS)、NAVSIM v2(90.3 EPDMS)和Navhard(35.0 EPDMS)上的实验展示了具有竞争力的驾驶性能。在navtest和内部高风险场景上的定性比较进一步说明了风险感知推理和危害响应轨迹生成。本文代码将公开发布于:此https URL。
英文摘要
Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git