发表机构
National Key Laboratory for Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University; TermiTech; School of Intelligence Science and Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室; 南京大学人工智能学院; TermiTech公司; 南京大学智能科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉语言导航(VLN)的高数据与资源开销问题,提出LookStep框架,在相同训练设置下于R2R-CE Val-Unseen任务取得49.7%成功率,且内存效率更高、数据使用更少。
AI 中文摘要
视觉语言导航(VLN)要求具智能体在未知环境中遵循自然语言指令。近期进展多由多模态大语言模型(MLLMs)推动。现有方法遵循下一步动作预测范式,仅监督专家动作,训练需大量数据,且依赖认知地图、累积历史帧或外部3D工具维护状态,导致高计算与内存开销。为实现资源高效的VLN,我们提出LookStep,一个统一端到端框架,结合以语言为中心的未来状态建模与事件驱动滚动记忆,利用语言标签为每个候选动作生成粗粒度导航进度与未来状态,同时自主决定是否将每个观测写入带语义角色的有界滚动记忆。我们通过实验验证LookStep,在VLN-CE任务中,LookStep在相同训练设置下优于现有方法,在R2R-CE Val-Unseen上达到49.7%的成功率,且具有更好的内存效率与更少的数据使用。代码和模型可在此https URL获取。
英文摘要
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing methods follow a next-step action prediction paradigm, supervising only the expert action, which requires a high quantity of data for training. They also rely on cognitive maps, accumulated historical frames, or external 3D tools to maintain states, leading to high computational and memory overhead. To realize resource efficiency VLN, we propose LookStep, a unified end-to-end framework that combines Language Centric Future State Modeling and Event Driven Rolling Memory that uses language labels to generate coarse-grained navigation progress and future states for each candidate action, while autonomously deciding whether to write each observation into a bounded rolling memory with a semantic role. We validate LookStep empirically. On VLN-CE tasks, LookStep outperforms existing methods under the same training settings, achieving a 49.7\% success rate on R2R-CE Val-Unseen with better memory efficiency and less data usage. Code and model is available at https://github.com/kunyang-YU/LookStep.
Comments19 Pages, 7 Figures. Accepted in EMNLP 2026 Main. Project Page: https://kunyang-yu.github.io/LookStep/