arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38641cs.CVcs.RO

基于语言记忆的视觉-语言-行动自动驾驶智能体

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AD-Memo,一种基于语言记忆的VLA自动驾驶智能体,通过将记忆作为思维链扩展输出,并采用SFT和半闭环RL算法训练,在全向停车场和一般驾驶场景中提升驾驶质量、问答能力,并提供可移植的记忆。

中文摘要 AI 辅助

视觉-语言-行动(VLA)基础模型最近已成为自动驾驶的主流解决方案之一,因为它们可以利用视觉-语言预训练期间获得的知识进行准确且可解释的驾驶。然而,由于图像的高令牌成本,VLA只能将有限数量的帧作为视觉输入,这对于依赖记忆的任务(如确定全向停车场的到达顺序和长视距驾驶场景理解)是有问题的。现有解决方案使用通过交叉注意力访问的潜在向量记忆,这些记忆既不可解释也不可移植。在本文中,我们提出了AD-Memo,一种具有基于语言记忆的通用VLA驾驶智能体。该智能体将其记忆作为思维链(CoT)的扩展输出,以记录对驾驶至关重要的周围物体;该记忆成为智能体未来输入的一部分。我们整理了基于记忆的数据集,并采用两阶段方案训练VLA:监督微调(SFT)和Da Capo,一种新颖的半闭环强化学习(RL)算法,该算法使用轨迹级优势用于记忆,步级优势用于驾驶,从而实现更好的信用分配。在全向停车场和一般驾驶等场景中,AD-Memo提高了驾驶质量,实现了对驾驶场景更好的问答,并为其他模型提供了即插即用的记忆。

英文摘要

Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑