arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31814cs.AI

DriveHierarchy:从开环理解到闭环执行的VLM驾驶能力诊断基准

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun

首次发表
浏览论文内容

中文总结 AI 辅助

DriveHierarchy是一个分层基准,将VLM驾驶能力分为四个等级,整合开环数据与闭环模拟,通过15个VLM实验验证其诊断和优化价值。

中文摘要 AI 辅助

评估基于视觉语言模型(VLM)的自动驾驶仍然困难,因为驾驶能力是复合性的,一个合格的系统必须能够理解交通参与者和危险,整合跨视角和时间的情境信息,推理未来的演变,并在闭环交互下采取适当行动。现有基准通常要么评估开环理解,要么评估闭环驾驶,但在解释这些能力如何组织、如何关联以及如何为模型诊断和改进提供信息方面,结构有限。我们提出DriveHierarchy,一个分层基准,将基于VLM的自动驾驶组织为四个等级,涵盖感知基础、情境记忆、心智推理和闭环执行。为了实例化这一层级结构,我们将多个开源自动驾驶数据集整合为一个统一的开环基准,包含84,279帧上的76,798个问答对,并开发了一个基于真实世界道路网络的闭环模拟平台,支持交互式场景构建,从中精选100个驾驶场景用于具身评估。对15个VLM的实验表明,DriveHierarchy捕捉了结构化但非冗余的能力变化,将开环理解与闭环驾驶联系起来,并为诊断和基准引导的优化提供了实用基础。因此,DriveHierarchy作为评估和改进基于VLM的自动驾驶系统的统一框架。一个匿名项目已在此https URL上发布。

英文摘要

Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on https://github.com/PerfectXu88/DriveHierarchy

补充信息

↑