arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PRISM:面向端到端自动驾驶运动规划的特权概率潜在监督

PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning

Volodymyr Havrylov, Faris Janjoš, Andreas Look, Jürgen Mathes, Andreas Geiger

arXiv 2608.01201首次发表:更新:

发表机构

University of Tübingen; Bosch Center for Artificial Intelligence; Coburg University; Tübingen AI Center(蒂宾根大学; 博世人工智能中心; 科堡大学; 蒂宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对端到端自动驾驶模型梯度微弱问题,提出基于真实值的概率潜在监督框架,在 nuScenes 数据集上实现规划 L2 误差降 8%、碰撞率降 3%,且计算开销可忽略。

AI 中文摘要

端到端自动驾驶(E2E AD)系统将感知、预测和规划整合为单一可微架构,尽管这类模型展现出巨大潜力,但其标准训练常仅依赖输出级监督,这会导致日益复杂模型的隐藏层梯度微弱。近期研究将视觉语言模型(VLM)监督用于潜在特征以解决该问题,取得了显著实证收益,但 underlying 理论机制却鲜为人知。我们对该方法的研究表明,性能提升并非源于 VLM 的推理能力,而是源于训练期间 E2E AD 模型与真实值(GT)数据之间建立的潜在连接。基于此见解,我们提出一种概率深度监督框架,直接从 GT 数据对中间潜在表示进行正则化,通过将模型潜在视为可重参数化分布,利用证据下界(ELBO)优化架构。我们在 nuScenes 数据集上开展的评估表明,用未来 GT 路径监督与轨迹相关的潜在特征可持续提升规划性能;在使用相同训练数据和 E2E 架构的情况下,与具有竞争力的向量化基线相比,我们的方法将规划 L2 误差降低 8%,碰撞率降低 3%,同时仅产生可忽略的计算开销。

英文摘要

End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single differentiable architecture. While these models show great promise, their standard training often relies on output-only supervision, which can lead to weak gradients for the hidden layers of increasingly complex models. Recent works have integrated vision-language model (VLM) supervision for latent features to address this, yielding substantial empirical gains, yet leaving the underlying theoretical mechanisms poorly understood. Our investigation into this methodology reveals that the resulting performance gains stem not from VLM reasoning capabilities, as previously assumed, but rather from the latent connections forged between the E2E AD model and ground-truth (GT) data during training. Building on this insight, we propose a probabilistic deep supervision framework that regularizes intermediate latent representations directly from GT data. By treating model latents as reparameterizable distributions, we optimize the architecture via the Evidence Lower Bound (ELBO). Our evaluations conducted on the nuScenes dataset demonstrate that supervising trajectory-related latents with future GT paths consistently improves planning performance. Using identical training data and E2E architectures, our method achieves an 8% reduction in planning L2 error and a 3% decrease in collision rates compared to competitive vectorized baselines, all while incurring negligible computational overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑