发表机构
Renmin University of China; Hebei Key Laboratory of Real-virtual Integrated Autonomous Systems (RIAS)(中国人民大学; 河北省虚实融合自主系统重点实验室(RIAS))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对协同驾驶中VLM问答与规划未联合评估的问题,提出真实世界基准CoVLM-Bench及统一基线CoVLM-Drive,通过配对视图实现CDQA和CP,实验验证了其有效性和规划误差降低。
AI 中文摘要
视觉语言模型(VLM)在自动驾驶领域取得了显著进展,但其成功主要是在以自我为中心的(ego-centric)场景中研究的。基础设施侧观测提供了超越自车视野的视角,然而传统的协同驾驶系统通常将这些观测转换为几何表示,用于下游感知和规划。将这些视角直接纳入VLM,为改进协同场景理解和轨迹规划提供了机会。然而,在相同的真实世界车辆-基础设施场景中,问答和轨迹规划尚未被联合评估。我们提出了CoVLM-Bench,一个在车辆-基础设施配对场景上用于协同驾驶问答(CDQA)和协同规划(CP)的基准。CoVLM-Bench提供了基于场景的CDQA标注、作为辅助监督的三部分理由(rationales),以及从记录的自我运动(ego motion)中导出的未来轨迹目标。它包含2,196个配对帧,带有35,136条CDQA标注,而CP在3秒时间范围内预测六个路点(waypoints)。标注结合了模型辅助起草、基于记录的计算和人工验证。基于CoVLM-Bench,我们引入了CoVLM-Drive,一个直接使用配对视图进行CDQA和CP的统一VLM基线。实验表明,CDQA适应提高了答案准确性,且CoVLM-Drive达到了比所比较的V2X规划器更低的最终位移误差(FDE);QA初始化和理由监督各自减少了规划误差。总之,CoVLM-Bench和CoVLM-Drive支持用于协同场景理解和规划的VLM的训练与比较。
英文摘要
Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperative planning (CP) on vehicle-infrastructure paired scenes. CoVLM-Bench provides scene-grounded CDQA annotations, three-part rationales as auxiliary supervision, and future trajectory targets derived from recorded ego motion. It contains 2,196 paired frames with 35,136 CDQA annotations, while CP predicts six waypoints over a three-second horizon. The annotations combine model-assisted drafting, record-based computation, and human verification. Built upon CoVLM-Bench, we introduce CoVLM-Drive, a unified VLM baseline that directly uses paired views for both CDQA and CP. Experiments show that CDQA adaptation improves answer accuracy and that CoVLM-Drive reaches a lower FDE than the compared V2X planners; QA initialization and rationale supervision each reduce planning error. Together, CoVLM-Bench and CoVLM-Drive support the training and comparison of VLMs for cooperative scene understanding and planning.
Comments32 pages, 11 figures, 13 tables. Main text 9 pages; appendices from page 17