发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对零样本目标导航方法在现实实时场景下的性能下降问题,提出RTNav架构,在HM3D系列数据集的实时变体上实现了成功率与完成时间加权成功率的显著提升。
AI 中文摘要
借助强大的视觉与语言基础模型,在未知环境中寻找未预见物体的导航任务已愈发可行,但这些模型也会引入不可忽视的推理延迟,当智能体必须在现实世界中持续运行时,延迟便成为重要考量。多数最先进的方法仍在同步模拟器中开发,这类模拟器会等待智能体行动,推理时间几乎无限制,因此智能体常围绕感知、推理、行动的顺序执行设计,几乎不考虑时间约束。在现实执行中,挂钟时间计入任务预算,这些架构的低效性便会凸显。我们发现,近期的零样本目标导航方法在这种现实时间条件下会出现持续的性能下降。基于此,我们提出RTNav,一种简单却有效的架构,将推理延迟、异步环境步进、受限计算作为明确的设计考量。在HM3D-v1、HM3D-v2和HM3D-OVON的实时变体上评估显示,RTNav相比现有工作,最高提升11%的成功率,以及最高5.1分的完成时间加权成功率。
英文摘要
Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference time is effectively free. As a result, agents are often designed around the sequential execution of perception, reasoning, and action, with little regard for time constraints. Under real-time execution, where wall-clock time counts towards the task budget, the inefficiencies of these architectures become clear. We show that recent zero-shot object navigation methods suffer consistent performance degradation under such realistic timing conditions. Motivated by this observation, we propose RTNav, a simple but effective architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design considerations. Evaluated on real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, RTNav improves the success rate by up to 11% and the Success weighted by Completion Time by up to 5.1 points over prior work.