发表机构
Uber AV Labs; Case Western Reserve University(优步自动驾驶实验室; 凯斯西储大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出 GeoWAM,一种用于自动驾驶的视觉几何世界动作模型,通过预训练预测未来场景几何来学习动力学,经评估其生成的驾驶策略比图像基方案更强,确立未来几何预测为自动驾驶有效预训练目标。
AI 中文摘要
世界动作模型(WAMs)作为一种在自动驾驶中联合建模场景演化与 ego 动作的框架,近来受到越来越多的关注。大多数现有 WAMs 通过结合用于未来观测预测的视频生成骨干网络与用于 ego 轨迹预测的动作头,在像素空间中学习场景动力学。然而,像素仅提供这些动力学的间接表示:它们将几何、运动与外观、纹理、光照纠缠在一起,迫使模型从二维观测中推断三维变换。我们认为,以点云表示的几何为驾驶提供了更自然的状态空间,因为它明确捕获了空间结构以及支配场景演化的刚性与非刚性变换,同时直接与执行驾驶动作的空间对齐。基于这一见解,我们提出 GeoWAM,一种面向自动驾驶的视觉几何世界动作模型。GeoWAM 不预测未来图像,而是被预训练以预测未来场景几何,生成的表示联合编码空间结构与时间演化。随后,几何条件动作头利用这些学习到的几何动力学来预测未来 ego 轨迹。大量开环与闭环评估表明,视觉几何世界建模比基于图像的替代方案产生显著更强的驾驶策略,确立未来几何预测为自动驾驶的有效预训练目标。
英文摘要
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving a combined EPDMS of 36.6 on navhard without PDMS supervision and strong zero-shot generalization to nuScenes, with a collision rate of 0.24%. Scaling geometry pretraining with unlabeled data further improves performance, increasing the navhard score by 8.2% to 39.6 and strengthening zero-shot transfer to nuScenes, where the collision rate is reduced by 50% to 0.12%. Together, these results establish geometry as an effective state representation for autonomous driving and geometry pretraining as a general, scalable strategy for downstream planning.
CommentsProject page: https://yiren-lu.com/project_pages/geowam/