arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09695cs.ROcs.CV

更好的视觉表征是否总能带来更好的端到端自动驾驶?

Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

  • Southeast University(东南大学)
  • OpenDriveLab at The University of Hong Kong(香港大学OpenDriveLab)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

Zihao Zhang, Haochen Tian, Tianyu Li, Changhui Jing, Jingliang He, Naisheng Ye, Ziyuan Pu, Zhenjie Yang

AI总结:

本研究提出ViRA框架探究视觉基础模型表征对端到端自动驾驶的影响,发现VFM表征普遍提升性能,但目标选择重要,辅助监督可降低敏感性,并据此开发ViRA-Diffusion规划器,在NAVSIM上达到92.3 EPDMS,领先至少1.9分。

AI中文摘要:

视觉基础模型(VFMs)因其强大的表征能力而日益被整合到端到端自动驾驶中,然而目前尚不清楚这些表征在何时能提升驾驶性能。为探究这一问题,我们提出了ViRA,一个与规划器无关的视觉表征对齐框架,该框架保持规划器架构和推理成本不变。我们的研究揭示了三个发现:(1)VFM引导的视觉表征在多种端到端规划器中持续提升驾驶性能,且增益扩展到零样本闭环评估。(2)VFM目标的选择对规划性能至关重要,与不同VFM的对齐可进一步惠益已采用预训练VFM编码器的规划器。(3)辅助感知监督降低了对VFM目标选择的敏感性,将五个目标间的EPDMS差异从2.7分缩小至0.5分,并可能补偿效果较弱的VFM目标。基于这些发现,我们开发了ViRA-Diffusion,一种基于扩散的规划器,在无辅助感知监督下训练,在NAVSIM v2 navtest上达到92.3 EPDMS,在我们比较中的最新方法中至少领先1.9分。这些结果促使在将VFM整合到端到端自动驾驶时,应联合考虑目标选择与规划器监督。结果和演示可访问此https URL获取。

英文摘要:

Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at https://github.com/OpenDriveLab/ViRA.

↑