发表机构
The University of Tokyo; TIER IV, Inc.; Huawei(东京大学; TIER IV公司; 华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型用于闭环驾驶时推理延迟与车辆控制速率冲突的问题,提出快慢架构,慢系统用冻结主干处理历史,快专家基于缓存和当前帧回归航路点,经训练在CARLA上提升路线完成率,降低误差且可零样本转移。
AI 中文摘要
大语言模型为端到端驾驶带来了指令跟随和场景推理能力,但其推理延迟与车辆所需的控制速率相冲突。现有闭环智能体通过在交替的模拟时间步调用模型并在其间重放先前命令来掩盖这一差距,导致一半的控制输出忽略了最新观测。我们提出了一种快慢架构来消除这种折衷。一个冻结的7B视觉-语言主干作为慢系统,低频消化导航指令和视觉历史,同时将其每层的键值缓存作为场景的固定表示。一个轻量级动作专家作为快系统,在每个模拟时间步关注此缓存和当前相机帧,通过单次前向传递回归航路点。由于缓存与实际场景存在延迟,我们在随机陈旧性下训练专家,使训练与异步执行对齐。在CARLA的LangAuto-Short路线上,我们的系统每50毫秒模拟时间步产生新的控制,将路线完成率从37.0提升到94.0,超过了跳帧基线。具有相同专家的跳帧消融实验分离了两个起作用的因素:专家自身提高了驾驶分数,而每时间步的新鲜度将完成率从82.提升到94.0,并将闯红灯违规减少了三分之一。在单个城镇训练的专家可以零样本转移到两个未见城镇,保持84-94%的路线完成率,而基线仅为31-41%。与主干自身的动作头相比,它将开环航路点误差降低了近四倍,在单个消费级GPU上每时间步模型成本为32毫秒,且与历史长度无关。
英文摘要
Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command in between, so half of all control outputs ignore the newest observations. We present a fast-slow architecture that removes this compromise. A frozen 7B vision-language backbone acts as the slow system, digesting navigation instructions and visual history at low frequency while exposing its per-layer key-value cache as a standing representation of the scene. A lightweight action expert acts as the fast system, attending to this cache and to the current camera frame at every simulation tick to regress waypoints in a single forward pass. Since the cache lags behind the world at deployment, we train the expert under randomized staleness, aligning training with asynchronous execution. On LangAuto-Short routes in CARLA, our system produces fresh control at every 50 ms simulation tick and lifts route completion from 37.0 to 94.0 over the frame-skipping baseline. A frame-skip ablation with the same expert separates the two factors at work: the expert raises the driving score on its own, while per-tick freshness raises completion from 82.1 to 94.0 and cuts red-light violations by a third. Trained on a single town, the expert transfers zero-shot to two unseen towns, holding 84-94% route completion where the baseline reaches 31-41%. It reduces open-loop waypoint error by nearly a factor of four compared to the backbone's own action head, at a per-tick model cost of 32 ms that is independent of history length on a single consumer GPU.
Comments13 pages, 5 figures, 4 tables