arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReDrive:通过世界建模塑造表征以实现端到端驾驶

ReDrive: Shaping Representations with World Modeling for End-to-End Driving

Yueting Zhu, Shaoyu Chen, Yuehao Song, Hui Sun, Qian Zhang, Wenyu Liu, Xinggang Wang

arXiv 2609.33854首次发表:更新:

发表机构

Huazhong University of Science & Technology; Horizon Robotics(华中科技大学; 地平线机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReDrive通过世界建模塑造表征,实现无需推理时辅助模块的高性能端到端驾驶规划,在NAVSIM上取得领先结果。

AI 中文摘要

驾驶策略需要具备场景理解和未来演化预测的能力。为实现这一目标,当前的端到端模型通常构建复杂的感知-规划流水线,或引入显式预测未来状态的世界模型,从而导致系统架构复杂。受通用视觉表征可迁移性的启发,我们认为,将足够强的视觉表征与表征世界建模相结合,可以在不依赖复杂推理时辅助模块的情况下支持有效的规划。基于这一见解,我们提出了ReDrive,一种通过未来表征预测来强化面向规划的视觉特征的端到端驾驶框架。为实现这一点,ReDrive采用三阶段训练流水线,包括驾驶视频预训练、联合世界建模与规划训练以及规划器适配。这产生了强面向规划的表征和高性能规划器,同时既不需要辅助感知模块,也不需要在推理时进行未来预测。在NAVSIM上的实验展示了强劲性能,在NAVSIM v1上达到91.0 PDMS,在NAVSIM v2上达到90.8 EPDMS。这些结果表明,用世界建模塑造表征足以实现高性能的端到端规划,同时保持简单的编码器-规划器推理流水线。

英文摘要

Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by the transferability of general-purpose visual representations, we argue that combining sufficiently strong visual representations with representation world modeling can support effective planning without relying on complex inference-time auxiliary modules. Based on this insight, we present ReDrive, an end-to-end driving framework that strengthens planning-oriented visual features via future representation prediction. To achieve this, ReDrive adopts a three-stage training pipeline consisting of driving video pretraining, joint world-modeling and planning training, and planner adaptation. This yields a strong planning-oriented representation and a high-performance planner, while requiring neither auxiliary perception modules nor future prediction at inference time. Experiments on NAVSIM demonstrate strong performance, achieving 91.0 PDMS on NAVSIM v1 and 90.8 EPDMS on NAVSIM v2. These results show that shaping representations with world modeling is sufficient to enable high-performance end-to-end planning while retaining a simple encoder-planner inference pipeline.

Comments15 pages,7 figures,10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑