arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先寻后动:视觉语言导航中进度锚定的证据搜寻

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang

arXiv 2609.37353首次发表:更新:

发表机构

Tsinghua University; Pengcheng Laboratory; South China University of Technology; International Digital Economy Academy; Ping An Technology (Shenzhen) Co., Ltd.; Harbin Institute of Technology (Shenzhen)(清华大学; 鹏城实验室; 华南理工大学; 国际数字经济学院; 平安科技(深圳)有限公司; 哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言导航中智能体因证据不足而盲目行动的问题,提出SeekVLN框架,通过未来引导反向生成和反事实对比策略优化训练,主动搜寻任务相关证据,在R2R-CE和RxR-CE上成功率分别提升12.7%和7.5%。

AI 中文摘要

视觉语言导航(VLN)要求智能体从长程指令和不完整的自我中心观察中持续锚定任务进度。现有的基于视觉语言模型(VLM)的导航智能体通常仅对可用观察进行推理,即使在缺少任务相关证据时也可能保持自信。例如,智能体可能自信地向前行进而迷路,即使指示下一次转弯的地标位于其当前视野之外。我们将这种失败模式称为“进度近视”(Progress Myopia):智能体未能识别不可靠的进度锚定,并继续基于不充分的证据行动。为解决此问题,我们提出了SeekVLN,一个证据搜寻框架,将语义进度推理与主动获取任务相关观察相结合。SeekVLN分两个阶段训练:首先,未来引导反向生成(FRG)利用未来专家动作,以补充视图和证据标注增强离线专家轨迹。在这些轨迹上的监督微调建立了证据搜寻和进度推理的先验,无需额外的专家交互。然而,仅靠模仿并不能揭示搜寻是否能改善后续导航。因此,我们引入了反事实对比策略优化(C2PO)进行强化微调。通过将每个证据搜寻分支与来自同一状态的反事实直接导航分支进行比较,C2PO使用对比奖励根据后续导航收益为搜寻决策分配信用。在模拟基准上的实验表明,SeekVLN实现了最先进的性能,在R2R-CE和RxR-CE上分别比基础模型提高了12.7%和7.5%的成功率。模拟和真实世界评估均展现出类似人类的证据搜寻行为,以实现更可靠的进度锚定。

英文摘要

Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑