用于自动驾驶视觉语言模型可验证推理的未来轨迹延迟暴露
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
浏览论文内容
中文总结 AI 辅助
该研究针对自动驾驶VLA模型的轨迹锚定偏差问题,提出AD-MCQ规划框架与DEFT-RLVR方法,将未来轨迹转化为决策后验证目标,提升了自动驾驶推理能力并保留了通用视觉能力。
中文摘要 AI 辅助
近期,自动驾驶领域的视觉-语言-动作(VLA)模型日益利用思维链(CoT)监督来提升其视觉语言模型(VLM)组件的推理能力,但现有标注流程通常会让教师模型接触到记录的真实未来轨迹(GT)。我们通过实证研究表明,这会引发轨迹锚定偏差:教师模型会为已揭示的结果进行合理化解释,而非从场景证据中推断决策,从而产生因果忠实度较低的思维链,且会出现更为严重的幻觉,尤其在因果关系复杂的场景中。移除真实轨迹可消除这种捷径,但开放式轨迹生成会将高层决策与精确的几何合成、低层动力学纠缠在一起。为在无需开放式轨迹合成的情况下实现轨迹级驾驶决策的可验证性,我们提出自动驾驶多项选择题(AD-MCQ),将规划转化为在明确轨迹候选中进行选择。在此基础上,我们提出用于强化学习与价值推理的未来轨迹延迟暴露方法(DEFT-RLVR),将未来轨迹从决策前的锚点转化为决策后的验证目标。实验结果显示,DEFT-RLVR可提升自动驾驶推理能力,同时保留甚至增强通用视觉能力。仅使用视觉语言模型推理,且可通过候选构建控制难度,AD-MCQ为未来可验证自动驾驶推理研究提供了灵活、可扩展且可拓展的基础。
英文摘要
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.