arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系锚定的自动驾驶潜在世界模型

Relationally Grounded Latent World Models for Autonomous Driving

Fabian Schmidt, Markus Enzweiler, Abhinav Valada

arXiv 2609.24626首次发表:更新:

发表机构

Esslingen University of Applied Sciences; University of Freiburg(埃斯林根应用科学大学; 弗莱堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用交通场景图作为语义监督,对齐潜在世界模型表示,在nuScenes上显著降低轨迹误差和碰撞率,验证了显式关系结构对自动驾驶表示学习的价值。

AI 中文摘要

潜在世界模型学习用于自动驾驶的预测性表示,但这些状态所保留的关系语义往往仍是隐式的。我们研究了交通场景图是否可以作为潜在世界表示的特权语义监督。基于LAW,我们从nuScenes 3D标注构建以智能体为中心的场景图,使用冻结的文本嵌入模型编码其序列化的关系结构,并在训练期间将视觉潜在表示与该语义目标对齐。在推理时我们移除监督分支,因此它既不需要场景图也不需要3D标注,且不增加测试时计算。在nuScenes上,与重新训练的LAW基线相比,我们的方法将平均轨迹L2误差从0.661降至0.622(降低5.9%),碰撞率从0.456降至0.217(降低52.4%)。它还优于非结构化的字幕式语义目标,支持显式关系结构对潜在世界模型表示学习的有益性。

英文摘要

Latent world models learn predictive representations for autonomous driving, but the relational semantics these states preserve often remain implicit. We investigate whether traffic scene graphs can serve as privileged semantic supervision for latent world representations. Building on LAW, we construct actor-centric scene graphs from nuScenes 3D annotations, encode their serialized relational structure using a frozen text embedding model, and align the visual latent representations with this semantic target during training. We remove the supervision branch at inference, so it requires neither scene graphs nor 3D annotations and adds no test-time computation. On nuScenes, our method reduces average trajectory L2 error from 0.661 to 0.622 (5.9%) and collision rate from 0.456 to 0.217 (52.4%) relative to our retrained LAW baseline. It also outperforms an unstructured caption-style semantic target, supporting the benefit of explicit relational structure for latent world-model representation learning.

CommentsAccepted at the NeuRo-SymBolic World Models (RoBoWoMo) Workshop at IROS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑