arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CMU-Drive与V2V-VLA:具备推理能力的协同多智能体统一驾驶基准及车对车视觉-语言-动作模型

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

Hsu-kuang Chiu, Stephen F. Smith

arXiv 2608.07621首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出协同多智能体统一驾驶基准CMU-Drive及车对车视觉-语言-动作模型V2V-VLA,解决现有VLA模型缺乏协同能力的问题,构建了协同自动驾驶研究的首个基准并将相关资源开源。

AI 中文摘要

视觉-语言-动作(VLA)模型近期在端到端自动驾驶任务中取得了出色性能,但现有方法主要针对单个自动驾驶智能体设计,对协同感知、推理及规划的支持有限。本文提出CMU-Drive(Cooperative Multi-agent Unified Driving with Reasoning),这是一种闭环端到端基准,用于评估多辆联网自动驾驶车辆(CAV)在涉及背景交通参与者的安全关键驾驶场景下的协同自动驾驶能力。我们进一步提出V2V-VLA(Vehicle-to-Vehicle Vision-Language-Action),这是一种协同VLA模型,通过联合生成驾驶动作、未来路径点、语言推理及通信策略,将协同驾驶整合至单次前向传播中。在CMU-Drive上开展的实验构建了首个协同VLA驾驶基准及基线,为多智能体、闭环、端到端协同自动驾驶的未来研究奠定了基础。本文代码、基准及模型检查点将公开发布,以促进开源研究。

英文摘要

Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑