发表机构
Amap, Alibaba Group; Beihang University(高德软件有限公司,阿里巴巴集团; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有视觉语言导航方法缺乏闭环机制问题,提出ReflectVLN框架,通过双向交互智能体决策,引入行动思维链训练方案,实验证明该框架在有限数据下提升成功率和路径效率,且具良好训练成本与可解释性。
AI 中文摘要
现有视觉语言导航方法常将视觉语言模型(VLM)与航点解码器结合生成多步行动计划,但缺乏明确闭环机制来跟踪语义进展、诊断执行失败及从长期导航中的错误积累中恢复。为填补这一空白,我们提出ReflectVLN,一个通过双向交互意图和执行智能体组织决策的智能体视觉语言导航框架。意图智能体执行子任务分解和反思,生成可执行的子任务描述作为纠正计划。执行智能体依据这些描述在当前观察下将其转化为短期行动,同时监测子目标进展并检测偏离行为。关键的是,ReflectVLN实现了闭环双向通信。为鼓励具有可解释中间推理的时间连贯决策,我们引入行动思维链(Action-CoT),一种用于行动生成的路径条件双查询训练方案。在标准视觉语言导航基准测试中的实验表明,ReflectVLN在有限数据预算下提高了成功率和路径效率,具有良好的训练成本,推理时高级意图调用更少,同时提供可解释的中间决策用于分析和协作。
英文摘要
Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN