arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28296cs.ROcs.HC

Talk2Escape:视觉与语言导航中的对话基础

Talk2Escape: Conversational Grounding for Vision-and-Language Navigation

  • Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所)
  • Zhejiang Wanli University(浙江万里学院)
  • Fudan University(复旦大学)
  • Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

Zerui Li, Sihao Lin, Yanyan Shao, Jiwen Zhang, Xiangyu Shi, Shijie Li, Qi Wu

AI总结:

Talk2Escape提出一种主动对话干预框架,将视觉与语言导航从开环转为闭环,通过轻量模块监测轨迹偏差并生成查询以获取纠正反馈,在模拟和真实环境中显著提升导航鲁棒性。

AI中文摘要:

尽管视觉与语言导航(VLN)已展现出显著的成功,但当前的单轮范式暴露了一个根本性弱点:智能体以严格的开环方式运行。在实践中,感知混淆、传感器噪声和里程计漂移等因素会导致微小偏差随时间累积,往往引发灾难性的任务失败,且缺乏内置的错误恢复机制。为解决这一问题,我们提出了Talk2Escape,一个主动且与模型无关的对话干预框架,将导航重新构建为闭环交互过程。其核心是一个轻量级视觉-语言模块,持续监测智能体的运动学状态。当检测到局部循环或严重的轨迹偏离时,该模块将原始自我中心观察转化为简洁且有根据的查询,以从算法预言机或人机回环中寻求针对性的纠正反馈。在包括R2R-CE、RxR-CE和VLNVerse在内的高保真模拟器中的广泛评估表明,Talk2Escape在不同基础智能体上均展现出一致的改进。实验上,Talk2Escape在R2R-CE上实现了66.0%的成功率,超越了当前有监督和零样本的最先进方法。我们进一步在Unitree Go2四足机器人上验证了其从仿真到现实的迁移,证明主动对话显著提高了物理环境中的导航鲁棒性。

英文摘要:

While Vision-and-Language Navigation (VLN) has demonstrated remarkable success, the prevailing single-turn paradigm exposes a fundamental vulnerability: agents operate in a strictly open-loop manner. In practice, factors such as perceptual aliasing, sensor noise, and odometry drift can cause minor deviations to accumulate over time, often leading to catastrophic mission failures with no built-in mechanism for error recovery. To address this, we introduce \textit{Talk2Escape}, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process. At its core, a lightweight vision-language module continuously monitors agent kinematics. Upon detecting localized looping or severe trajectory divergence, it translates raw egocentric observations into concise, grounded queries to solicit targeted corrective feedback from either an algorithmic oracle or a human-in-the-loop. Extensive evaluations in high-fidelity simulators, including R2R-CE, RxR-CE, and VLNVerse, demonstrate that \textit{Talk2Escape} exhibits consistent improvements across diverse base agents. Empirically, \textit{Talk2Escape} achieves a 66.0\% Success Rate on R2R-CE, outperforming the current supervised and zero-shot state-of-the-art methods. We further validate its sim-to-real transfer on a Unitree Go2 quadruped, proving that proactive dialogue drastically improves navigation robustness in physical environments.

补充信息

↑