AI 中文总结
针对现有V2X基准的局限,推出V2XBench平台与Chat-V2XBench数据集,提出AURORA端到端协同驾驶框架,其在严重遮挡场景下路径完成率达98.21%、驾驶得分76.02,开创了可扩展的V2X-VLM范式。
AI 中文摘要
车万物互联(V2X)协同支持非视距感知,可缓解单车感知中的遮挡问题。但现有V2X基准对闭环评估和语言 grounding 监督的支持有限,阻碍了用于端到端协同驾驶的视觉-语言模型(VLM)的发展。为解决这些局限,我们推出V2XBench,这是一个具备同步自车-路侧感知和闭环评估的仿真平台,同时推出Chat-V2XBench,一个用于协同推理的渐进式结构化视觉问答(VQA)数据集。基于该基准基础设施,我们提出AURORA,一个端到端协同驾驶框架。AURORA配备双视图感知架构,通过查询级跨视图查询对齐与融合(CQAF)模块缓解自车与路侧视角间的空间和语义差异。利用生成的统一 token,经LoRA适配的VLM连接语义推理与生成式轨迹规划。在V2XBench上开展的大量闭环评估显示,AURORA在严重遮挡场景中达到了最先进性能,路径完成率为98.21%,驾驶得分为76.02,同时所需路侧通信带宽较低。最终,本研究开创了可扩展的V2X-VLM范式,为下一代协同自动驾驶铺平道路。
英文摘要
Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-end cooperative driving. To address these limitations, we introduce V2XBench, a simulation platform featuring synchronized ego--roadside sensing and closed-loop evaluation, together with Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning. Building upon this benchmark infrastructure, we propose AURORA, an end-to-end cooperative driving framework. Equipped with a dual-view perception architecture, AURORA mitigates spatial and semantic discrepancies across ego and roadside viewpoints through a query-level Cross-View Query Alignment and Fusion (CQAF) module. Leveraging the resulting unified tokens, a LoRA-adapted VLM bridges semantic reasoning and generative trajectory planning. Extensive closed-loop evaluations on V2XBench demonstrate that AURORA achieves state-of-the-art performance in heavily occluded scenarios, with a Route Completion rate of 98.21% and a Driving Score of 76.02, while requiring low roadside communication bandwidth. Ultimately, this work pioneers an extensible V2X--VLM paradigm, paving the way for next-generation cooperative autonomous driving.
Comments15 pages, 5 figures, under review