发表机构
Institute of Automation, Chinese Academy of Sciences; Chongqing Chang’an Technology Co., Ltd.(中国科学院自动化研究所; 重庆长安科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DriveVLA-M0,一种具备故障感知潜在记忆的检索增强型VLA模型,通过故障案例记忆与针对性修正机制,在NAVSIM基准上优于现有方法,实现自动驾驶性能提升。
AI 中文摘要
视觉-语言-动作(VLA)模型近期成为端到端自动驾驶领域极具潜力的范式,可实现感知、语言与规划的统一推理。但现有方法缺乏利用过往故障或适应分布偏移的机制,导致模型在曾发生故障的相似场景中持续表现不佳。本文提出DriveVLA-M0,一种具备故障感知潜在记忆的检索增强型VLA模型。我们构建了一个潜在记忆池,存储故障案例及其场景结构表征与专家轨迹标签,并设计了专用检索模型,解耦静态道路结构与动态智能体交互,以实现基于结构的检索。推理阶段,检索到的案例通过轻量级解耦LoRA的测试时训练(TTT)机制注入模型,无需修改主干网络即可实现针对性的场景特定修正。在NAVSIMv1与NAVSIMv2基准上的大量实验表明,我们的方法始终优于现有方法,在Navtest上达到94.1 PDMS、Navhard上达到47.0 EPDMS,且仅产生26.44 ms的TTT反向延迟开销。此外,我们证明DriveVLA-M0可通过增加内存有效扩展,无需额外训练即可通过内存扩展获得性能提升。代码可在指定URL获取。
英文摘要
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However, existing approaches lack mechanisms to exploit past failures or adapt to distribution shifts, causing the model to persistently underperform on similar scenarios where it has previously failed. In this paper, we propose DriveVLA-M0, a retrieval-augmented VLA with failure-aware latent memory. We construct a latent memory pool that stores failure cases along with their structure scene representations and expert trajectory labels, and design a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval. At inference time, retrieved cases are injected into the model via a lightweight decoupled LoRA-based test-time training (TTT) mechanism, allowing targeted and scenario-specific correction without modifying the backbone. Extensive experiments on NAVSIMv1 and NAVSIMv2 benchmark demonstrate that our approach consistently outperforms prior methods, achieving 94.1 PDMS on Navtest and 47.0 EPDMS on Navhard with only 26.44 ms TTT backward latency overhead. Furthermore, we show that DriveVLA-M0 scales effectively with additional memory, enabling training-free performance gains through memory expansion. The code is available at https://github.com/ZebinX/DriveVLA-M0.