发表机构
University of California, Los Angeles; NVIDIA; dot by Hyundai; Fudan University(加利福尼亚大学洛杉矶分校; 英伟达; 现代旗下42dot公司; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出 AffordDrive3D 模型,基于 VLM 骨干联合建模未来动作相关区域与几何,在 NAVSIM 数据集上取得 91.3 PDMS、89.9 EPDMS 的 SOTA 性能,提升自动驾驶轨迹规划效果。
AI 中文摘要
世界-动作模型近期通过联合学习未来场景预测与轨迹生成,提升了自动驾驶性能。现有多数方法主要通过 RGB 外观对未来进行建模,近期研究开始引入几何预测以改善空间理解。然而,密集几何仅描述了整个场景的空间布局,未指明哪些部分与 ego 车辆的动作最相关。对于驾驶而言,模型还必须识别并预判自身可安全行驶的区域,以及可能存在碰撞风险的区域。联合建模与动作相关的区域和未来几何,可为策略提供驾驶相关线索及其对应的空间结构。因此,我们提出 AffordDrive3D,这是一种具备 affordance 和几何感知的世界-动作模型,可联合学习与未来动作相关的区域和空间结构。为捕捉驾驶 affordance 预测所需的场景语义和驾驶上下文,我们基于 VLM 骨干构建 AffordDrive3D,以预测直接影响 ego 运动的可行驶区域和碰撞关键区域,同时从 RGB 世界模型潜变量中预测未来几何。在 NAVSIM 数据集上,AffordDrive3D 实现了 91.3 PDMS 和 89.9 EPDMS 的 SOTA 性能,证明了联合建模未来 affordance 和几何对轨迹规划的有效性。
英文摘要
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.