arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01755cs.AI

用于自动驾驶视觉语言模型可验证推理的未来轨迹延迟暴露

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Xiaozhi Chen, Yikun Ban, Deqing Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对自动驾驶VLA模型的轨迹锚定偏差问题,提出AD-MCQ规划框架与DEFT-RLVR方法,将未来轨迹转化为决策后验证目标,提升了自动驾驶推理能力并保留了通用视觉能力。

中文摘要 AI 辅助

近期,自动驾驶领域的视觉-语言-动作(VLA)模型日益利用思维链(CoT)监督来提升其视觉语言模型(VLM)组件的推理能力,但现有标注流程通常会让教师模型接触到记录的真实未来轨迹(GT)。我们通过实证研究表明,这会引发轨迹锚定偏差:教师模型会为已揭示的结果进行合理化解释,而非从场景证据中推断决策,从而产生因果忠实度较低的思维链,且会出现更为严重的幻觉,尤其在因果关系复杂的场景中。移除真实轨迹可消除这种捷径,但开放式轨迹生成会将高层决策与精确的几何合成、低层动力学纠缠在一起。为在无需开放式轨迹合成的情况下实现轨迹级驾驶决策的可验证性,我们提出自动驾驶多项选择题(AD-MCQ),将规划转化为在明确轨迹候选中进行选择。在此基础上,我们提出用于强化学习与价值推理的未来轨迹延迟暴露方法(DEFT-RLVR),将未来轨迹从决策前的锚点转化为决策后的验证目标。实验结果显示,DEFT-RLVR可提升自动驾驶推理能力,同时保留甚至增强通用视觉能力。仅使用视觉语言模型推理,且可通过候选构建控制难度,AD-MCQ为未来可验证自动驾驶推理研究提供了灵活、可扩展且可拓展的基础。

英文摘要

Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.

↑