DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
DriveWorld-VLA: 基于视觉-语言-动作的统一潜在空间世界建模用于自动驾驶
机构 * School of Computer Science(计算机科学学院) ; Beijing Key Laboratory of Traffic Data Mining(交通数据挖掘重点实验室) ; Beijing Jiaotong University(北京交通大学)
专题命中 端到端驾驶 :autonomous driving(title,abstract);分类 cs.RO、cs.CV
AI总结 DriveWorld-VLA通过统一视觉-语言-动作与潜在空间世界建模,提升自动驾驶中的决策与前瞻性想象能力。
Comments 20 pages, 7 tables, 12 figures