发表机构
Zhejiang University; Huawei Technologies Co., Ltd; Zhejiang Humanoid Robot Innovation Center(浙江大学; 华为技术有限公司; 浙江人形机器人创新中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对连续环境视觉语言导航的模仿学习缺陷,提出分层马尔可夫决策过程结合拓扑图与动作感知价值头的方法,在R2R-CE等基准取得最优性能。
AI 中文摘要
连续环境中的视觉语言导航(VLN-CE)要求智能体遵循自然语言指令在未见过的环境中移动。现有模仿学习(IL)流程在这种闭环设置中表现不佳:行为克隆存在分布偏移问题,而DAgger的专家动作在轨迹偏离时会变得模糊。虽然强化学习(RL)是解决该问题的自然范式,但由于奖励稀疏,直接将RL应用于微观动作空间样本效率低下。为克服这一瓶颈,我们将VLN-CE重新表述为分层马尔可夫决策过程(MDP),明确将高层规划与低层控制解耦。通过将环境抽象为拓扑图,我们的高层策略在前沿节点构成的宏观动作空间上运行,无需训练的低层控制器作为其状态转移,这显著压缩了决策时域,使闭环RL变得可行。为支持宏观MDP上的RL优化,我们提出一种动作感知价值头,以有效评估动态前沿动作空间下的状态价值,为基于图的PPO提供支持。大量实验证明了我们架构的有效性,最终我们的模型在R2R-CE和RxR-CE基准上达到了最先进的性能。
英文摘要
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language instructions through unseen environments. Existing imitation learning (IL) pipelines struggle in this closed-loop setting: behavior cloning suffers from distribution shift, and DAgger's expert actions become ambiguous upon trajectory deviation. While Reinforcement Learning (RL) offers a natural paradigm to address this, directly applying RL to micro action spaces is sample-inefficient due to reward sparsity. To overcome this bottleneck, we reformulate VLN-CE as a Hierarchical Markov Decision Process (MDP), explicitly decoupling high-level planning from low-level control. By abstracting the environment into a topological graph, our high-level policy operates on a macro action space of frontier nodes, with a training-free low-level controller acting as its state transition, which significantly compresses the decision horizon and makes closed-loop RL tractable. To support RL optimization on the macro MDP, we propose an action-aware value head to effectively evaluate state values under the dynamic frontier action space, powering a graph-based PPO. Extensive experiments demonstrate the effectiveness of our architecture. Finally, our model achieves state-of-the-art performance on the R2R-CE and RxR-CE benchmarks.