发表机构
University of Michigan; Tsinghua University(密歇根大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人在连续场景中从室内到室外的导航问题,引入NavVerse基准,涵盖多种场景和任务。通过可执行接口用多指标评估智能体,零样本实验显示当前智能体在跨上下文导航上差距大,适应是主要瓶颈。
AI 中文摘要
部署在配送、校园和应急响应场景中的机器人通常需要在单个连续场景中从建筑物导航到街道。现有基准通常分别评估室内和室外导航,许多基准忽略了机器人执行,使得出口寻找、边界穿越、适应和运动动力学故障未得到充分探索。我们引入了NavVerse,这是一个用于室内到室外具身导航的启用物理的基准。NavVerse包含100个室内场景、50个城市室外场景和50个室内到室外场景,以及跨越对象导航、视觉与语言导航和地点导航任务的10000个场景,其中智能体搜索餐厅或银行等语义兴趣点。通过可执行机器人接口使用任务成功、路径效率和安全指标对智能体进行评估。使用RL、VLA和模块化基线的零样本实验表明,当前智能体在解决跨上下文导航方面仍有很大差距:端到端VLA获得最高的零样本成功率,而模块化方法提供最强的安全配置文件。地点导航进一步揭示了从室外到室内到室外场景的明显下降,表明适应仍然是主要瓶颈。
英文摘要
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored. We introduce NavVerse, a physics-enabled benchmark for indoor-to-outdoor embodied navigation. NavVerse contains 100 indoor scenes, 50 urban outdoor scenes, and 50 indoor-to-outdoor scenes, and 10,000 episodes spanning Object Navigation, Vision-and-Language Navigation, and Place Navigation tasks, where agents search for semantic points of interest such as restaurants or banks. Agents are evaluated through executable robot interfaces using task-success, path-efficiency, and safety metrics. Zero-shot experiments with RL, VLA, and modular baselines show that current agents remain far from solving cross-context navigation: end-to-end VLAs obtain the highest zero-shot success, while the modular method provides the strongest safety profile. PlaceNav further reveals a clear drop from outdoor to indoor-to-outdoor scenes, indicating that adaptation remains major bottleneck.