AI 中文总结
DrivingBench是首个要求视觉语言模型驾驶真实汽车的基准,通过控制丰田卡罗拉绕锥桶赛道,测试了GPT-6 Astra等模型在推理延迟下的驾驶能力,Astra是唯一完成赛道的模型。
AI 中文摘要
前沿模型在许多数字基准测试中表现出色,然而它们驾驶真实汽车的能力——这一日常人类技能——在很大程度上仍未得到测试。我们提出了DrivingBench,据我们所知,这是第一个要求通用视觉语言模型驾驶真实汽车的基准测试。通过三种工具,模型能够看到来自丰田卡罗拉的摄像头画面,并直接控制其转向和速度,在低速下绕停车场锥桶赛道行驶。当模型思考时,汽车可以继续移动,新的命令会替换当前正在执行的命令,因此推理延迟成为任务的一部分,测试模型在这种约束下观察、行动、监控、恢复以及完成长时程目标的能力。我们在供应商原生的工具链(Codex、Claude Code、Cursor)中对GPT-6 Astra、Claude Fable 5.1、GPT-5.6 Sol和Grok 4.6进行了基准测试,每个模型在一次对话中最多尝试三次;Astra是唯一完成赛道的模型,它在第二次尝试中完成,而其他模型的任何尝试均未通过赛道的50%。四个模型中有两个在保留上下文的情况下跨尝试取得了实质性改进。我们还详细介绍了我们动作接口背后的设计原则,并展示了工具输出格式和任务框架如何共同决定模型是会驾驶还是拒绝。我们发布了我们的工具链、提示词、赛道地图以及带有视频和遥测数据的轨迹,以供复现。
英文摘要
Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We also detail the design principles behind our action interface, and show how the tool output format and the framing of the task combined to determine whether models would drive at all or refuse. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.