DriveZero:超越人类示范的端到端驾驶
DriveZero: End-to-End Driving Beyond Human Demonstrations
- Xiaomi EV(小米汽车)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DriveZero通过分离感知与动作模型,结合DriveRL闭环强化学习和DriveVFM多基础模型融合,实现超越人类示范的端到端驾驶,在nuPlan和NAVSIM等基准上达到最先进性能。
AI中文摘要:
大多数端到端自动驾驶系统通过模仿人类驾驶日志进行学习,导致其学习到的行为受限于记录轨迹的质量和行为覆盖范围。本报告提出了DriveZero,一个学习超越人类示范的驾驶行为的端到端系统。它将驾驶分解为感知模型和动作模型,在各自最适合的领域分别进行预训练,并将它们组合成一个端到端规划器。这两个模型需要不同的学习方案:感知必须理解世界,因此受益于大规模且多样化的视觉数据;动作必须与世界交互,因此需要闭环反馈。在动作方面,我们引入了DriveRL,一个混合智能体闭环强化学习框架。它将真实的驾驶日志转换为交互式世界,在其中通过闭环轨迹滚动,使用PPO算法训练一个特权教师策略。对于感知模型,DriveVFM将多个冻结的视觉基础模型(包括DINOv3、SigLIP2、SAM和Depth Anything V2)整合到一个仅基于原始图像的单骨干网络中,无需任务特定的标注。DriveZero随后将两者统一:一个仅使用摄像头的规划器,通过其轨迹滚动来蒸馏冻结的DriveRL教师。此外,目标条件的教师可以在增强的驾驶意图下进行查询,从而产生日志数据无法提供的多样化、目标一致的监督。在nuPlan上,带有价值引导的测试时动作搜索的DriveRL在Val14、Test14-hard和Test14-random社区划分中,在非反应式和反应式模式下均取得了93.57的平均分数,在所有三个划分上均超过了Log-Replay专家。DriveZero在NAVSIMv1、NAVSIMv2和闭环HUGSIM基准上取得了最先进的性能,且无需任何人类轨迹监督。
英文摘要:
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.