arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06055cs.CV

DriveZero:超越人类示范的端到端驾驶

DriveZero: End-to-End Driving Beyond Human Demonstrations

  • Xiaomi EV(小米汽车)

机构由 AI 辅助整理,请以论文原文为准。

Hao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang,… 展开作者

Hao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang

AI总结:

DriveZero通过分离感知与动作模型,结合DriveRL闭环强化学习和DriveVFM多基础模型融合,实现超越人类示范的端到端驾驶,在nuPlan和NAVSIM等基准上达到最先进性能。

AI中文摘要:

大多数端到端自动驾驶系统通过模仿人类驾驶日志进行学习,导致其学习到的行为受限于记录轨迹的质量和行为覆盖范围。本报告提出了DriveZero,一个学习超越人类示范的驾驶行为的端到端系统。它将驾驶分解为感知模型和动作模型,在各自最适合的领域分别进行预训练,并将它们组合成一个端到端规划器。这两个模型需要不同的学习方案:感知必须理解世界,因此受益于大规模且多样化的视觉数据;动作必须与世界交互,因此需要闭环反馈。在动作方面,我们引入了DriveRL,一个混合智能体闭环强化学习框架。它将真实的驾驶日志转换为交互式世界,在其中通过闭环轨迹滚动,使用PPO算法训练一个特权教师策略。对于感知模型,DriveVFM将多个冻结的视觉基础模型(包括DINOv3、SigLIP2、SAM和Depth Anything V2)整合到一个仅基于原始图像的单骨干网络中,无需任务特定的标注。DriveZero随后将两者统一:一个仅使用摄像头的规划器,通过其轨迹滚动来蒸馏冻结的DriveRL教师。此外,目标条件的教师可以在增强的驾驶意图下进行查询,从而产生日志数据无法提供的多样化、目标一致的监督。在nuPlan上,带有价值引导的测试时动作搜索的DriveRL在Val14、Test14-hard和Test14-random社区划分中,在非反应式和反应式模式下均取得了93.57的平均分数,在所有三个划分上均超过了Log-Replay专家。DriveZero在NAVSIMv1、NAVSIMv2和闭环HUGSIM基准上取得了最先进的性能,且无需任何人类轨迹监督。

英文摘要:

Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.

↑