发表机构
Joy Future Academy(京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉与语言导航问题,提出JOP-VLN框架,通过三阶段训练流程,协同结合离线策略模仿学习和在线策略探索,采用高熵轨迹采样和纠错优先轨迹排序策略,在相关基准上取得高成功率,创技术新水平。
AI 中文摘要
视觉与语言导航(VLN)要求具身智能体依据自然语言指令在物理世界中导航。视觉语言模型(VLM)的进展推动了基于VLM的VLN方法发展,主要有两种范式:基于专家示范的模仿学习(IL)及数据集聚合(DAgger)算法增强错误恢复能力;基于可验证奖励的强化学习(RL)提升推理与探索能力。本文介绍JOP-VLN,一种在三阶段训练流程中协同结合离线策略模仿学习和在线策略探索的新框架。先通过IL获取基本导航技能,再用DAgger生成启发式探索轨迹用于模仿学习改进错误恢复,最后实施联合在线与离线策略学习框架,采用高熵轨迹采样提高RL训练效率,用纠错优先轨迹排序策略有效纠错。实验表明JOP-VLN有效,在VLN-CE R2R和RxR基准上分别达到69.9%和68.0%的成功率,在R2R上创最新技术水平。
英文摘要
Vision-and-Language Navigation (VLN) necessitates an embodied agent to navigate in the physical world by adhering to natural language instructions. Recent advancements in Vision-Language Models (VLM) have propelled the development of VLM-based VLN methods with two predominant paradigms: (1) imitation learning (IL) on expert demonstrations, followed by the Dataset Aggregation (DAgger) algorithm to bolster error recovery capabilities; (2) reinforcement learning (RL) driven by verifiable rewards to enhance reasoning and exploration. A notable gap is the absence of integration between these two distinct paradigms. This paper introduces JOP-VLN, a novel VLN framework that synergistically combines off-policy imitation learning and on-policy exploration within a three-stage training pipeline. Initially, IL is employed on expert demonstrations to acquire basic navigation skills. Subsequently, the DAgger algorithm is utilized to generate heuristic exploration trajectories, which are then used for imitation learning to improve error recovery capabilities. Finally, a joint on-and-off policy learning framework is implemented, featuring high-entropy trajectory sampling to enhance RL training efficiency and an error-correction-prioritized trajectory sorting strategy for effective error correction. Extensive experiments demonstrate the efficacy of JOP-VLN, achieving success rates of 69.9% and 68.0% on the VLN-CE R2R and RxR benchmarks, respectively, setting a new state-of-the-art on R2R. Project page: https://qingrongh.github.io/JOP-VLN.
CommentsAccepted by IROS 2026