发表机构
Stony Brook University; Rutgers University; Sunrise Technology Inc.(石溪大学; 罗格斯大学; 晨昇科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出结果引导蒸馏的师生框架,将VLM的显式推理与几何轨迹生成结合,在Waymo基准上使自动驾驶性能提升约24%,优于经典推理基线。
AI 中文摘要
端到端(E2E)自动驾驶旨在学习从视觉观测到控制动作的直接映射,但这类E2E模型常作为黑箱运行,在复杂场景中表现不佳。为解决该问题,近期研究引入视觉语言模型(VLM)以提供显式推理,提升可解释性与驾驶鲁棒性。不过这些方法通常依赖预生成标注,存在标签可能出错、人力成本高昂的缺陷。本研究提出一种新框架,通过师生架构整合结构化推理与几何精度:教师模型引入反思性推理,VLM生成逻辑解释后,在真实动作监督下反思优化推理,无需中间标签即可提升零样本泛化能力;学生模型通过监督微调蒸馏教师的推理能力,还设计了独立的 waypoint 解码器,将文本推理转化为连续轨迹。该方案兼顾两大目标:提供显式推理以增强可解释性,同时实现鲁棒准确的驾驶性能,通过阶段推理引擎利用二者协同提升驾驶表现,且明确用推理引导驾驶预测。在Waymo基准上的评估显示,本框架在零样本推理、waypoint准确率和推理效率上均优于经典基于推理的基线;实验验证了该设计,推理文本对驾驶推理有显著贡献,与不含推理的相同模型相比,性能提升约24%。本研究推动推理驱动的自动驾驶向可解释、可部署系统发展。
英文摘要
End-to-end (E2E) autonomous driving aims to learn a direct mapping from visual observations to control actions. However, these E2E models often act as black boxes and struggle with complex scenarios. To address this, recent works incorporate Vision-Language Models (VLMs) to provide explicit reasoning, enhancing both interpretability and driving robustness. These approaches typically rely on pre-generated annotations, which suffer from potentially flawed labels and require costly human labor. In this work, we propose a new framework that integrates structured reasoning and geometric precision through a teacher-student architecture. The teacher model introduces reflective reasoning, where the VLM generates logical explanations and then reflectively refines the reasoning under the supervision of ground-truth action. This enhances zero-shot generalization without intermediate labels. The student model distills the teacher's reasoning capabilities via supervised fine-tuning. We also design a separate waypoint decoder that interprets textual reasoning into continuous trajectories. Our proposed solution integrates two goals: providing explicit reasoning for interpretability and delivering robust and accurate driving performance. It leverages the synergy between these two objectives within a staged inference engine to enhance driving performance and explicitly uses the reasoning to guide driving prediction. Evaluated on Waymo benchmarks, our framework outperforms classical reasoning-based baselines in zero-shot reasoning, waypoint accuracy, and inference efficiency. Our experiments validate this design, demonstrating that the reasoning text makes a significant contribution to driving inference, resulting in around a 24% improvement in performance compared to an identical model that lacks reasoning. Our work advances reasoning-driven autonomous driving toward interpretable and deployable systems.