arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30818cs.ROcs.AI

评估即一切:多模态自动驾驶

Evaluation Is All You Need for Multi-Modal Autonomous Driving

Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, … 展开作者

Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出iDriveVLA多模态规划框架,通过统一轨迹评估器与渐进训练策略解决生成-评估不对称,在NAVSIM v1上达到94.95 PDMS,超越人类专家。

中文摘要 AI 辅助

多模态规划在自动驾驶中具有前景,通过在模糊和长尾场景中表示多种合理行为。现有方法主要集中于改进轨迹多模态性、增强轨迹表示或重塑候选分布。然而,我们识别出多模态规划中一个显著的生成-评估不对称性:尽管具有强大的oracle性能,现有规划器往往无法可靠地选择最佳可用候选,留下大量未实现的规划潜力。为解决这一挑战,我们提出iDriveVLA,一种多模态规划框架,改进候选轨迹空间,同时实现更可靠和上下文感知的轨迹评估。具体而言,iDriveVLA引入统一轨迹评估器,包含用于质量和风险估计的安全感知评分器,以及用于场景自适应标准加权的VLM引导调制器。我们进一步开发了oracle对齐的渐进训练策略,包括候选模仿预训练、候选空间细化和语义排名对齐。在公共NAVSIM v1排行榜上,iDriveVLA达到94.95 PDMS的新最先进性能,超越人类专家参考。

英文摘要

Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, leaving substantial planning potential unrealized. To address this challenge, we propose iDriveVLA, a multi-modal planning framework that improves the candidate trajectory space while enabling more reliable and context-aware trajectory evaluation. Specifically, iDriveVLA introduces a unified trajectory evaluator comprising a Safety-aware Scorer for quality and risk estimation, together with a VLM-guided Modulator for scene-adaptive criterion weighting. We further develop an oracle-aligned progressive training strategy consisting of candidate imitation pretraining, candidate space refinement, and semantic ranking alignment. On the public NAVSIM v1 leaderboard, iDriveVLA achieves a new state-of-the-art performance of 94.95 PDMS, surpassing the human-expert reference.

发表机构

  • School of Vehicle and Mobility & College of AI, Tsinghua University(清华大学车辆与运载学院与人工智能学院)
  • Dongfeng Motor Corporation Research & Development Institute(东风汽车集团有限公司研发总院)

机构由 AI 辅助整理,请以论文原文为准。

↑