arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Orbis 2:一种用于驾驶的分层世界模型

Orbis 2: A Hierarchical World Model for Driving

Sudhanshu Mittal, Arian Mousakhan, Silvio Galesso, Karim Farid, Jonannes Dienert, Rajat Sahay, Thomas Brox

arXiv 2607.15898首次发表:更新:

发表机构

University of Freiburg(弗莱堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对当前世界模型不足提出分层驾驶世界模型,跨两层分解预测,结合扩散强制预训练和教师强制微调,在长距离生成保真度等多方面取得领先成果。

AI 中文摘要

当前的世界模型在单一抽象层次上运行,大多数注重感知保真度,却缺乏现实世界下游任务所需的空间推理和语义理解。我们提出了一种分层驾驶世界模型,它在不同的时间和抽象尺度上跨两个层次对未来预测进行分解:一个高层次预测器,预测长时间范围内的粗略场景结构;一个低层次生成器,根据高层次输出生成详细预测。这种分解在产生高感知保真度的同时,还捕捉到强大的空间和语义表示。我们进一步表明,与标准的教师强制目标相比,使用扩散强制目标进行预训练会产生更丰富的内部表示,而教师强制(仅根据干净的上下文预测下一帧)会产生更稳定的自回归展开。因此,我们引入了一种通用的两阶段训练范式,先用扩散强制对模型进行预训练,再用教师强制进行微调,将前者的表示优势与后者的展开稳定性结合起来。我们的方法在既定基准上的标准驾驶世界模型评估套件中取得了领先成果,包括长距离生成保真度、在反事实场景中评估的转向响应性以及内部表示质量。项目页面提供代码、演示、检查点和定性结果:此https URL

英文摘要

Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world model that factorizes future prediction across two levels operating at distinct temporal and abstraction scales: a high-level predictor that forecasts coarse scene structure over extended temporal horizons, and a low-level generator that produces detailed predictions conditioned on the high-level output. This decomposition yields high perceptual fidelity while also capturing strong spatial and semantic representations. We further show that pretraining with a diffusion forcing objective yields substantially richer internal representations than the standard teacher forcing objective, while teacher forcing -- predicting only the next frame from clean context -- produces more stable autoregressive rollouts. We therefore introduce a generic two-stage training paradigm that pretrains the model with diffusion forcing and fine-tunes with teacher forcing, combining the representational benefits of the former with the rollout stability of the latter. Our approach achieves state-of-the-art results across the standard suite of driving world model evaluations on established benchmarks, including long-horizon generation fidelity, steering responsiveness evaluated on counterfactual scenarios, and internal representation quality. Project page with code, demo, checkpoints and qualitative results: https://lmb-freiburg.github.io/orbis2.github.io/

CommentsProject page: https://lmb-freiburg.github.io/orbis2.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑