AI 中文总结
本文提出LifelongCrossNav框架,结合持久3D语义记忆与跨楼层可通行性建模,在自研HM3D-MFMON基准上实现更优的多层顺序多目标导航性能。
AI 中文摘要
目标驱动导航在语义感知与探索方面已取得显著进展,但多目标导航与跨楼层导航的持久记忆通常被分开处理。本文提出LifelongCrossNav,一种面向未知多层室内环境中顺序多目标目标导航的框架。在每个 episode 中,智能体接收有序的目标查询序列,同时持续维护共享的稀疏3D语义体素记忆。该记忆逐步积累几何结构、可通行性状态及视觉-语言特征,使后续目标查询可检索已获取的场景信息,无需重建地图。为支持跨楼层的持久搜索,LifelongCrossNav结合了支持感知的3D可通行性建图、楼梯专用感知及方向感知的楼梯通行策略。统一导航策略协调同楼层前沿探索、实时与历史兴趣点检索、楼梯导航及目标对象的搜索与接近。本文还引入基于HM3D场景构建的顺序多楼层多目标导航基准HM3D-MFMON,包含一个专用子集,完成全部目标子任务序列至少需要一次楼层转换。实验结果表明,在HM3D-MFMON上,LifelongCrossNav的性能始终优于代表性的平面持久语义地图基线,证明持久3D语义记忆与跨楼层可通行性建模可有效支持多层环境中的顺序多目标导航。项目页面:this https URL
英文摘要
Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately. We present LifelongCrossNav, a framework for sequential multi-object ObjectNav in unknown multi-floor indoor environments. Within each episode, the agent receives an ordered sequence of object-goal queries while continuously maintaining a shared sparse 3D semantic voxel memory. This memory incrementally accumulates geometric structure, traversability states, and vision-language features, allowing subsequent object-goal queries to retrieve previously acquired scene information without rebuilding the map. To support persistent search across floors, LifelongCrossNav combines support-aware 3D traversability mapping, stair-specific perception, and direction-aware stair traversal. A unified navigation policy coordinates same-floor frontier exploration, live and historical point-of-interest retrieval, stair navigation, and target-object search and approach. We further introduce HM3D-MFMON, a benchmark for sequential Multi-Floor Multi-Object Navigation built on HM3D scenes, including a dedicated subset in which completing the full sequence of object-goal subtasks requires at least one floor transition. Experimental results show that LifelongCrossNav consistently outperforms a representative planar persistent semantic-map baseline on HM3D-MFMON, demonstrating that persistent 3D semantic memory and cross-floor traversability modeling effectively support sequential multi-object navigation in multi-floor environments. Project page: https://flageval-baai.github.io/LifelongCrossNavPage.