发表机构
QwenTeam; Huazhong University of Science and Technology(阿里巴巴千问团队; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究团队提出Qwen-Drive-1.0,在保留通用视觉-语言能力的同时,将3D感知、视觉问答、运动规划集成于统一框架,经多类评估证实其3D感知与运动规划性能具竞争力。
AI 中文摘要
我们提出了Qwen-Drive-1.0,这是面向自动驾驶的视觉-语言基础模型的初步探索。Qwen-Drive-1.0保留了预训练视觉-语言模型(VLM)的架构,并在统一框架内集成了3D感知、视觉问答与运动规划功能。一个外部鸟瞰图(BEV)感知头可联合执行3D目标检测、语义占据预测和BEV地图分割,它可作为共享表征所获取的3D信息的探测工具,还能为3D场景结构提供可检查的显式接口。规划专家模块以共享VLM表征为条件生成未来的自车轨迹。分阶段训练方案将驾驶监督与通用视觉-语言数据相结合,以获取驾驶特定能力,同时助力保留广泛的视觉理解和指令遵循能力。实验表明,该模型具备出色的3D感知和驾驶场景理解能力,同时在很大程度上保留了通用视觉-语言能力。在开环、伪闭环和闭环设置下的全面评估进一步显示其运动规划性能极具竞争力。
英文摘要
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
CommentsCode will be available at https://github.com/QwenLM/Qwen-Drive-1.0