arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少语言,多隐变量:面向驾驶的标注高效VLA

Less Language, More Latents: Annotation-Efficient VLAs for Driving

Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania

arXiv 2609.27747首次发表:更新:

发表机构

Robert Bosch GmbH; Five AI Ltd.(罗伯特·博世有限公司; Five AI 有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA训练中语言标注稀缺问题,提出LADA三阶段流水线,利用隐变量动作模型和少量标注实现高效语言条件驾驶,在Bench2Drive上以不到5%标注超越全监督基线。

AI 中文摘要

视觉-语言-动作模型(VLA)有望实现人类可操控的自动驾驶,但其训练受到与自然语言指令配对的帧稀缺性的瓶颈制约:虽然摄像头流和专家轨迹被大规模记录,但语言标注(例如“在交叉路口左转”)仍然稀缺且获取成本高昂。为解决这一挑战,我们引入了隐变量动作驾驶标注(LADA),这是一个三阶段流水线,将大量未标注的观测-轨迹对转化为语言条件控制的基底。首先,我们训练一个带有向量量化瓶颈的隐变量动作模型,生成一个紧凑的高层车辆意图码本。其次,使用一个小的语言标注子集训练一个视觉-语言翻译器,将观测和语言指令映射到该码本中。第三,我们在整个未标注语料库上,基于观测-隐变量动作对训练一个驾驶VLA。仅使用不到5%的语言标注,且不依赖任何辅助的思维链推理或视觉问答流,LADA在闭环Bench2Drive基准上达到了87.98的驾驶分数和70.46%的成功率,匹配或超越了全监督基线。

英文摘要

Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑