EvoNav-Bench:演化环境中的终身导航基准
EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments
浏览论文内容
中文总结 AI 辅助
提出EvoNav-Bench基准,在演化环境中扩展终身导航任务,通过引入环境修改评估现有方法,并比较三种启发式策略以分析智能体适应场景变化的能力。
中文摘要 AI 辅助
终身导航(LN)要求具身智能体在同一环境中解决一系列导航子任务。由于从头解决每个子任务会导致冗余探索,LN智能体必须整合早期阶段的经验并在后续阶段中复用,通常通过持久化场景表示(如场景图或视觉快照)来实现。然而,现有方法通常假设环境是静态的,而在现实世界的LN场景中,人类活动可能导致环境发生演化。当静态假设被违背时,现有方法可能会将过时的先验观测与新观测融合,但当前基准无法揭示这种失效模式。在本文中,我们提出了EvoNav-Bench,它在演化环境的背景下扩展了GOAT-Bench风格的LN公式。基于ProcTHOR框架构建,EvoNav-Bench在导航任务之间引入环境修改,使先验经验有用但并非完全可靠。这种设计能够对环境演化如何影响复用先验场景观测的LN智能体进行受控评估。利用EvoNav-Bench,我们对三种构建并复用场景表示进行导航的最新方法进行了基准测试。我们还比较了三种处理环境演化的简单启发式策略:Frontier-Update、Fail-then-Update和Stage-Reset。我们的结果表明,现有方法在环境演化下表现脆弱,而启发式策略则能够对智能体如何适应场景变化并减轻其影响进行受控分析。
英文摘要
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exploration, an LN agent must consolidate experience from earlier stages and reuse it in later stages, often through persistent scene representations such as scene graphs or visual snapshots. However, existing approaches typically assume a stationary environment, whereas in real-world LN settings, human activities can cause the environment to evolve. With the stationary assumption violated, existing methods may fuse outdated prior observations with new observations, yet current benchmarks cannot reveal this failure mode. In this paper, we present EvoNav-Bench, which extends the GOAT-Bench style LN formulation in the context of evolving environments. Built on the ProcTHOR framework, EvoNav-Bench introduces environment modifications between navigation tasks, making prior experience useful but not fully reliable. This design enables controlled evaluation of how environment evolution affects LN agents that reuse prior scene observations. Using EvoNav-Bench, we benchmark three recent methods that build and reuse scene representations for navigation. We also compare three simple heuristic strategies for handling environment evolution: Frontier-Update, Fail-then-Update, and Stage-Reset. Our results show that existing methods are brittle under environment evolution, while the heuristic strategies enable a controlled analysis of how agents can adapt to scene changes and mitigate their impact.
发表机构
- State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室(BIGAI))
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。