MemVLN:用于视觉与语言导航的情景记忆和程序记忆
MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- The Chinese University of Hong Kong(香港中文大学)
- LIGHTSPEED(光速)
- The University of Hong Kong(香港大学)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究连续环境中视觉与语言导航问题,提出MemVLN框架,利用视觉编码器、大语言模型,通过情景记忆管理和程序记忆,实现实时推理,提升导航成功率并降低推理延迟。
AI中文摘要:
连续环境中的视觉与语言导航(VLN-CE)要求智能体在低延迟执行动作时保持长距离视觉历史以确保轨迹一致性。现有基于视频的VLN方法通常难以同时满足这两个需求。为应对这些挑战,我们提出了MemVLN,这是一个新颖的VLN框架,能以实时推理效率(14 FPS)实现领先性能。MemVLN利用视觉编码器处理连续观察,并使用大语言模型解释指令和生成动作。我们的方法核心是采用金字塔分辨率的情景记忆管理,集中计算于即时感知,同时保留压缩的长期历史。此外,引入程序记忆以通过紧凑的原子中级动作词汇实现快速动作,绕过自回归解码延迟。在VLN-CE上的实验表明,MemVLN-4B在R2R中比基线Qwen3-VL-4B架构的成功率高出5.8%,在RxR中高出9.7%,同时推理延迟加快了7倍。
英文摘要:
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typically struggle to satisfy both demands simultaneously. To address these challenges, we propose MemVLN, a novel VLN framework that achieves state-of-the-art performance with real-time inference efficiency (14 FPS). MemVLN utilizes a visual encoder to process continuous observations and a Large Language Model (LLM) to interpret instructions and generate actions. Central to our approach is an Episodic Memory management that applies pyramidal resolutions. This mechanism concentrates computation on immediate percepts while retaining compressed long-term history. Complementing to this design, we introduce Procedural Memory for fast action with a compact vocabulary of atomic mid-level actions to bypass auto-regressive decoding latency. Experiments on VLN-CE show that MemVLN-4B surpasses the baseline Qwen3-VL-4B architecture by 5.8\% SR in R2R and 9.7\% SR in RxR, while achieving a 7$\times$ speedup in inference latency.