arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLN-AVP:用于自动代客泊车的基于混合长短时记忆的零样本视觉语言导航

VLN-AVP: Zero-Shot Vision-Language Navigation with Hybrid Long-Short-Term Memory for Autonomous Valet Parking

Yijian Li, Xiangru Mu, Changze Li, Hantian Shi, Jiyuan Cai, Jia Cai, Xiaoxue Liu, Yajing Sun, Ming Yang, Tong Qin

arXiv 2607.17767首次发表:更新:

发表机构

Shanghai Jiao Tong University; Yinwang Intelligent Technology Co., Ltd.(上海交通大学; 银望智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自主代客泊车依赖预建地图限制可扩展性问题,提出VLN-AVP零样本导航框架,结合BEV模型与VLM,引入混合记忆系统,构建数据集和基准,实验证明该方法在模拟和实际车辆实验中均有良好效果,提升了成功率。

AI 中文摘要

自主代客泊车(AVP)中的现有方法通常依赖预建地图,这严重限制了它们在未见环境和开放词汇目标上的可扩展性。受视觉语言模型(VLM)在视觉语言导航(VLN)任务中的应用启发,我们提出了VLN-AVP,一种用于AVP任务的零样本导航框架。通过将鸟瞰图(BEV)模型的精确空间感知与VLM的通用智能相结合,该框架消除了对预建地图的依赖,解释停车场景中的语义环境上下文,并能根据自然语言指令进行直观导航。具体而言,我们引入了一种混合记忆系统:短期感知记忆跟踪语义视觉线索以解决现有方法中VLM单帧推理的局限性,而长期拓扑记忆则促进从过去经验中进行稳定的策略学习。为了弥合现有基准之间的差距,我们还展示了VLN-AVP数据集和基准。它具有10个高保真停车场景和超过1000个导航情节,是迄今为止车库场景数量最多的,也是第一个用于地下停车的VLN基准。大量实验表明,在模拟中,我们的方法与VLN方法相比成功率提高了25%以上,与其他自动驾驶方法相比提高了15%以上。此外,它在实际车辆实验中获得了领先的成功率,证明了其实际可行性。

英文摘要

Existing methods in Autonomous Valet Parking (AVP) typically rely on pre-built maps, which severely restricts their scalability to unseen environments and open-vocabulary targets. Inspired by the application of Vision-Language Models (VLMs) in Vision-Language Navigation (VLN) tasks, we propose VLN-AVP, a zero-shot navigation framework for AVP tasks. By combining the precise spatial perception of a Bird's-Eye-View (BEV) model with the general intelligence of VLMs, our framework 1) eliminates the dependency on pre-built maps, 2) interprets semantic environmental contexts in parking scenarios, and 3) enables intuitive navigation following natural language instructions. Specifically, we introduce a hybrid memory system: a short-term perception memory tracks semantic visual cues to address the limitations of VLM's single-frame reasoning in existing methods, while a long-term topological memory facilitates stable policy learning from past experiences. To bridge the gap in existing benchmarks, we also present the VLN-AVP dataset and benchmark. Featuring 10 high-fidelity parking scenes and over 1,000 navigation episodes, it has the largest number of garage scenes to date and is the first VLN benchmark for underground parking. Extensive experiments demonstrate that in simulation, our method achieves an over 25% improvement in success rate compared to VLN methods and an over 15% improvement compared to other autonomous driving methods. Furthermore, it attains a leading success rate in real-world vehicle experiments, proving its practical feasibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑