RoXDrive:基于动作忠实回滚的端到端自动驾驶闭环强化学习
RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
查看机构详情
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- XPeng Motors(小鹏汽车)
- Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出RoXDrive,一种即插即用的闭环强化学习框架,通过动作-视觉忠实度评估器筛选世界模型回滚,实现端到端自动驾驶策略优化,在nuScenes和内部数据集上显著降低安全违规。
中文摘要 AI 辅助
端到端自动驾驶策略通常通过对记录的演示进行模仿学习来训练,而无需观察自身行为的后果,这导致在闭环真实世界部署中出现因果混淆。为解决此问题,强化学习(RL)后训练提供了一种有前景的替代方案,利用世界模型作为交互式训练环境,生成未来场景以改进策略。然而,现有方法要么依赖基于重建的模拟器,提供有限的反事实交互,要么采用合成模拟器实现长时域闭环交互,但代价是存在显著的模拟到现实差距。近期,视频世界模型展现出生成逼真多步未来回滚的能力,但可能无法忠实反映动作条件,导致动作-视觉不匹配。本文提出RoXDrive,一种即插即用的闭环强化学习框架,通过识别动作忠实的世界模型回滚实现可靠策略优化,包含两个阶段:1)模型预训练:除基于模仿的策略预训练外,我们设计了动作-视觉忠实度评估器,用于逆动力学估计,并辅以几何感知的辅助轨迹监督,从而实现对视觉动态是否忠实反映条件化自我车辆动作的长时域评估。2)动作忠实的强化学习后训练:智能体与世界模型迭代交互,形成长时域场景回滚,仅保留动作忠实的回滚用于密集的安全感知评分和场景级闭环强化学习后训练。在nuScenes及包含超过13万个训练场景的内部数据集上的大量实验表明,各规划器均获得一致改进,其中DiffusionDrive在nuScenes上将安全违规减少27.6%,Qwen3-VL在内部数据上减少33.7%。
英文摘要
End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.