GeoWorldAD:用于自动驾驶的几何世界行动模型
GeoWorldAD: Geometry World Action Model for Autonomous Driving
- Nanyang Technological University(南洋理工大学)
- Xiaomi EV(小米汽车)
- Zhejiang University(浙江大学)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究自动驾驶中安全高效规划决策问题,提出GeoWorldAD模型,通过在自我对齐3D空间中规划轨迹、用潜在未来几何标记预测场景演变,并逐步聚合多尺度几何线索,实验证明该模型在自动驾驶方面性能先进。
AI中文摘要:
自动驾驶需要在动态3D环境中做出安全且高效的规划决策。尽管近期的视觉/视频行动模型能直接从视觉观察中学习策略并随视觉Transformer和大规模训练数据发展良好,但常缺乏明确的几何基础和对未来的空间引导。本文提出GeoWorldAD,一种在自我对齐的3D空间中进行轨迹规划并通过潜在未来几何标记预测短视距场景演变的几何世界行动模型。当前几何为安全规划提供基本空间约束,未来几何揭示周围物体和以自我为中心的自由空间如何演变,减少过度保守决策且不牺牲安全性。通过迭代轨迹细化逐步聚合多尺度当前几何和潜在未来几何以有效利用这些几何线索。在NAVSIM v1和v2上的实验证明了其先进性能,凸显了明确的3D几何基础和未来几何世界建模对安全高效自动驾驶的有效性。
英文摘要:
Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.