发表机构
Southeast University; Universiti Malay; Ant Group; Tsinghua University; Peking University(东南大学; 马来亚大学; 蚂蚁集团; 清华大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有端到端自动驾驶方法在空间物理证据融合、推理适应及推理路径优化上的不足,提出FactorDrive框架,通过PCF-CoT数据集与QS-GRPO算法优化推理,在两类基准上取得最优规划性能。
AI 中文摘要
视觉语言模型(VLMs)已在场景理解方面取得进展,并支持端到端自动驾驶中的显式推理。然而,现有方法未能充分将空间物理证据融入规划推理,且推理适应方式较为粗糙,无法满足特定场景的规划需求。此外,自动驾驶后训练中,用于提升规划质量的推理路径优化问题在很大程度上尚未被探索。为解决这些局限,我们提出FactorDrive,一种由规划关键因素(PCFs)驱动的自适应多步推理端到端自动驾驶框架。我们首先开展大规模驾驶领域指令微调,以建立基础驾驶知识。在此基础上,我们构建PCF-CoT,一种思维链(CoT)数据集,该数据集将规划推理建立在与轨迹相关的空间物理证据之上,并围绕特定场景的PCFs组织推理,使推理路径的组合与深度能够适应不同的规划需求。我们进一步引入质量搜索引导的组相对策略优化(QS-GRPO),该方法通过轨迹级规划奖励引导蒙特卡洛树搜索(MCTS),以发现规划质量更高的推理路径,并利用得到的响应通过GRPO优化策略,从而提升轨迹规划性能。在开环(nuScenes)和面向闭环(NAVSIM)基准上开展的大量实验表明,FactorDrive实现了最优的规划性能。
英文摘要
Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.