arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAP-Nav:空间语义表示与主动感知结合的分层开放词汇对象导航

SAP-Nav: Spatial Semantic Representation Meets Active Perception for Hierarchical Open-Vocabulary Object Navigation

Xuetong Pei, Jian Liu, Vidura Munasinghe, Bo Miao, U-Xuan Tan, Wenrui Ding, Na Zhao

arXiv 2608.12707首次发表:更新:

AI 中文总结

该研究提出SAP-Nav框架,结合空间语义表示与主动感知,解决分层开放词汇对象导航问题,在基准测试中性能最优且真实场景可行,无需特定训练或预计算场景地图。

AI 中文摘要

分层开放词汇对象导航(OVON)要求智能体在未知环境中遵循自由形式指令,指令可能通过场景、房间、区域和实例级线索指定目标。尽管近期研究LangMap已将该场景形式化,但在部分观测下可靠解决该问题仍具挑战:空间定位需要持续的环境级证据,而目标验证需要清晰且具区分性的候选视角。我们提出SAP-Nav,这是一个完全在线、零样本的框架,通过主动感知满足这两项需求。SAP-Nav从主动获取的房间视角逐步构建可查询的空间语义表示,支持从任何已探索位置进行空间语义查询;还采用主动视角验证,评估当前观测是否提供足够证据,必要时重新定位智能体至更具信息性的视角,再依据类别和属性约束验证候选目标。SAP-Nav虽为分层OVON设计,但无需任务特定训练或预计算场景地图,即可支持分层和标准类别级OVON。在LangMap和HM3D-OVON上的实验显示,SAP-Nav实现了整体最佳性能,其中区域级导航的成功率(SR)较基于训练的方法提升12.2%;真实世界机器人实验进一步证明了其实际可行性,代码将在论文接收后公开。

英文摘要

Hierarchical open-vocabulary object navigation (OVON) requires agents to follow free-form instructions that may specify targets through scene-, room-, region-, and instance-level cues in unseen environments. Although recent work LangMap has formalized this setting, reliably solving it under partial observations remains challenging: spatial grounding requires persistent environment-level evidence, whereas target verification requires clear and discriminative candidate views. We present SAP-Nav, a fully online, zero-shot framework that addresses both requirements through active perception. SAP-Nav incrementally constructs a Queryable Spatial-Semantic Representation from actively acquired room views, enabling spatial semantic queries from any explored location. It further employs Active Viewpoint Verification to assess whether the current observation provides sufficient evidence and, when necessary, reposition the agent to a more informative viewpoint before verifying candidates against category and attribute constraints. Although designed for hierarchical OVON, SAP-Nav supports both hierarchical and standard category-level OVON without task-specific training or precomputed scene maps. Experiments on LangMap and HM3D-OVON show that SAP-Nav achieves the overall best performance, including a 12.2% improvement in SR over training-based methods on region-level navigation. Real-world robot experiments further demonstrate its practical feasibility. Code will be made publicly available upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑