基于语言记忆的视觉-语言-行动自动驾驶智能体
Vision-Language-Action Autonomous Driving Agent with Language-based Memory
浏览论文内容
中文总结 AI 辅助
本文提出AD-Memo,一种基于语言记忆的VLA自动驾驶智能体,通过将记忆作为思维链扩展输出,并采用SFT和半闭环RL算法训练,在全向停车场和一般驾驶场景中提升驾驶质量、问答能力,并提供可移植的记忆。
中文摘要 AI 辅助
视觉-语言-行动(VLA)基础模型最近已成为自动驾驶的主流解决方案之一,因为它们可以利用视觉-语言预训练期间获得的知识进行准确且可解释的驾驶。然而,由于图像的高令牌成本,VLA只能将有限数量的帧作为视觉输入,这对于依赖记忆的任务(如确定全向停车场的到达顺序和长视距驾驶场景理解)是有问题的。现有解决方案使用通过交叉注意力访问的潜在向量记忆,这些记忆既不可解释也不可移植。在本文中,我们提出了AD-Memo,一种具有基于语言记忆的通用VLA驾驶智能体。该智能体将其记忆作为思维链(CoT)的扩展输出,以记录对驾驶至关重要的周围物体;该记忆成为智能体未来输入的一部分。我们整理了基于记忆的数据集,并采用两阶段方案训练VLA:监督微调(SFT)和Da Capo,一种新颖的半闭环强化学习(RL)算法,该算法使用轨迹级优势用于记忆,步级优势用于驾驶,从而实现更好的信用分配。在全向停车场和一般驾驶等场景中,AD-Memo提高了驾驶质量,实现了对驾驶场景更好的问答,并为其他模型提供了即插即用的记忆。
英文摘要
Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。