具身视觉与语言导航的全面综述与系统真实世界评估
A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation
浏览论文内容
中文总结 AI 辅助
综述具身视觉与语言导航研究,按动作和模型范式组织方法并分析优缺点,在物理机器人平台对代表性配置进行真实世界评估,揭示模拟与真实部署的性能差距,强调未来研究在感知、决策和控制方面的关键挑战。
中文摘要 AI 辅助
导航是自主系统的基本能力,但现有方法大多依赖高度结构化模型和强先验假设,限制了其在开放不确定现实环境中的鲁棒性。视觉与语言导航(VLN)以数据驱动方式让机器人整合自然语言理解与视觉感知,提供了有前景的方向。本综述对VLN研究进行全面回顾,按动作范式(分层和整体框架)和模型范式(判别和生成方法)组织最先进方法并分析优缺点。还在物理机器人平台上对代表性VLN系统配置进行系统真实世界评估。十个不同真实场景实验表明测试配置下模拟与真实世界部署存在性能差距,如代表性整体仅RGB方法模拟成功率61%,真实世界降至22%,分层框架真实世界成功率51%。最后强调了未来研究中感知、决策和控制方面的关键挑战。
英文摘要
Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments. Vision-and-Language Navigation (VLN) offers a promising direction by enabling robots to integrate natural language understanding with visual perception in a data-driven manner. Although VLN has attracted increasing research attention, systematic methodological taxonomy and real-world validation remain limited. This survey presents a comprehensive review of VLN research. Specifically, state-of-the-art methods are organized along two orthogonal dimensions: action paradigms, including hierarchical and monolithic frameworks, and model paradigms, including discriminative and generative approaches. A critical analysis of their respective strengths and limitations is provided. Additionally, we conduct a systematic real-world evaluation of representative VLN system configurations on a physical robotic platform. Experiments across ten diverse real-world scenes show a substantial performance gap between simulation and real-world deployment under the tested configurations: a representative monolithic RGB-only method achieves 61% success in simulation but drops to 22% in real-world deployment, while a hierarchical framework achieves a higher real-world success rate of 51%, suggesting stronger robustness in our evaluation setting. Finally, we highlight key challenges in perception, decision-making, and control that must be addressed in future research.
发表机构
- College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。