发表机构
Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NavJev通过以动作中心的视觉压缩和判别性动作-语义记忆,将导航重构为紧凑视觉压缩加轻量动作选择,在R2R-CE上以低延迟实现高成功率。
AI 中文摘要
近年来,零样本视觉-语言导航(VLN)方法越来越依赖多模态大语言模型(MLLMs)来对视觉观察、导航指令和候选动作进行推理。尽管这些方法有效,但在每个导航步骤中反复调用自回归多模态推理会引入显著的推理延迟,限制了具身智能体的响应能力。我们提出了NavJev,一个高效的VLN框架,它将在线导航从重复的多模态生成重构为紧凑的视觉压缩,随后进行轻量级的类型化动作选择。具体而言,以动作中心的视觉压缩(ACVC)将航点几何、BLIP描述和RAM语义标签整合为候选动作的紧凑表示,而判别性动作-语义记忆(DASM)则过滤共享语义,并在导航步骤间维护判别性的动作特定证据。基于这些表示,Jev直接对可用动作集进行结构化的概率决策。在R2R-CE上的实验表明,NavJev实现了27.0%的成功率(SR)和22.4%的路径长度加权成功率(SPL),每个导航步骤仅需0.65秒,同时与基于MLLM的VLN方法相比,显著降低了推理延迟和成本。项目页面可在以下网址获取:此https URL。
英文摘要
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.