AI 中文总结
SUV是将未来场景理解建模为视频生成的端到端驾驶框架,其在NAVSIM-v2和WOD-E2E基准上优于近期SOTA方法,实现了更优的轨迹规划性能。
AI 中文摘要
端到端驾驶需要对未来场景进行连贯理解,但现有方法使用特定任务的头部和输出格式对这些场景进行建模,可扩展性有限。视频生成能否提供一个共享预测器?我们提出SUV,这是一个统一的端到端驾驶框架,利用预训练的视频基础模型将未来场景理解建模为视频生成。SUV通过共享视频专家将未来外观、语义、相对深度和实例级动态建模为视频流,无需特定流的视觉预测头部。通过联合视频-动作注意力,动作专家关注所有未来流的潜在表示并生成自车轨迹。实验表明,SUV直接预测所有四个未来流,而受控 ablation 显示结构化未来监督和直接未来流访问可产生更高的轨迹规划分数。仅使用单个前置摄像头且无候选轨迹选择的情况下,SUV在NAVSIM-v2的两个划分上均优于大量近期SOTA方法,在navtest上达到91.0 EPDMS,在navhard上达到36.9;在长尾WOD-E2E基准上,SUV达到7.94的竞争力RFS。
英文摘要
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.
Comments16 pages, 5 figures. Code: https://github.com/ASH-2046/SUV