发表机构
Robert Bosch GmbH; Five AI Ltd.(罗伯特·博世有限公司; Five AI 有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FIVE-VLA通过高效视觉编码器和循环动作记忆模块,以6.41亿参数实现快速有效的自动驾驶,在Bench2Drive上路线完成率提升10%,碰撞率显著降低,速度提升8-30倍。
AI 中文摘要
最先进的自动驾驶视觉-语言-动作模型(VLA)面临关键限制:参数数量过多、高分辨率图像处理效率低下以及缺乏时间记忆。我们提出了快速有效VLA(FIVE-VLA),通过两个关键贡献来解决这些问题。首先,我们采用高效的视觉编码器,处理高分辨率(448×896)图像时仅生成98个令牌,比现有方法少5倍以上,并完全绕过文本生成以实现单次轨迹预测。其次,我们提出了循环动作记忆(RAM),这是一个轻量级模块,根据先前的动作令牌来条件化动作预测,为超车和紧急制动等操作提供关键的时间上下文。仅用6.41亿参数,FIVE-VLA在具有挑战性的Bench2Drive闭环驾驶基准上,比之前最先进的VLA多完成约10%的路线且无交通规则违规。在大规模真实世界NVIDIA Physical AI AV数据集上的非反应式开环模拟显示,在单视图和四视图设置中,碰撞违规率分别比SimLingo低10.2%和7.7%。此外,FIVE-VLA在A100上运行约30帧/秒,在T4 GPU(边缘设备的代理)上运行约4帧/秒,比先前方法快8-30倍。
英文摘要
State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.