AI 中文总结
SparseNav提出免训练的指令条件稀疏语义感知框架,按需获取语义并维护轻量BEV地图和地标记忆,在R2R-CE和RxR-CE上分别取得42.8%和40.7%的成功率,并成功部署于四足机器人。
AI 中文摘要
基于地图的视觉语言导航(VLN)依赖持久化的空间表示来连接语言理解与几何规划。然而,获取超出当前指令需求的语义信息可能会引入不必要的感知成本和无关标注。持续累积无关物体不仅浪费计算资源,还会使视觉语言模型(VLM)规划器所消费的视觉-空间表示变得杂乱。为解决这一问题,我们提出了SparseNav,一个遵循“少即是多”原则进行语义导航的免训练框架。SparseNav持久化维护一个轻量级的几何鸟瞰图(BEV)地图和稀疏地标记忆,利用主动子指令按需获取新语义,以决定哪些内容值得进行接地(grounding)。指令管理器首先跟踪导航进度并识别活动地标查询。随后,指令条件感知机制在查询地标可见且其度量位置能为下一步决策提供信息时,调用开放词汇分割。由此产生的地标记忆支持VLM在混合前沿和局部方向航点候选之间进行选择。无需任何额外训练,SparseNav在R2R-CE和RxR-CE的Val-Unseen分割上分别实现了42.8%和40.7%的成功率。受控消融实验考察了语义感知策略和各个框架组件的贡献。此外,我们成功将SparseNav部署在配备Intel RealSense D455 RGB-D相机(用于几何建图和地标接地)以及Livox MID-360 LiDAR(用于定位)的Unitree Go2四足机器人上,无需预建地图。我们使用指令条件航点导航在多个室内环境中验证了其有效性。
英文摘要
Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unrelated objects may not only waste computation, but also clutter the visual-spatial representation consumed by the vision-language model (VLM) planner. To address this problem, we present SparseNav, a training-free framework that follows a less-is-more principle for semantic navigation. SparseNav persistently maintains a lightweight geometric bird's-eye-view (BEV) map and sparse landmark memory, acquiring new semantics on demand using the active sub-instruction to decide what is worth grounding. An instruction manager first tracks navigation progress and identifies the active landmark query. An instruction-conditioned perception mechanism then invokes open-vocabulary segmentation when the queried landmark is visible and its metric location can inform the next decision. The resulting landmark memory supports VLM selection among hybrid frontier and local directional waypoint candidates. Without any additional training, SparseNav achieves success rates of 42.8% on R2R-CE and 40.7% on RxR-CE, both on the Val-Unseen splits. Controlled ablations examine semantic perception strategies and the contributions of individual framework components. Furthermore, we successfully deployed SparseNav on a Unitree Go2 quadruped equipped with an Intel RealSense D455 RGB-D camera for geometric mapping and landmark grounding and a Livox MID-360 LiDAR for localization, without a prebuilt map. We validated its effectiveness across multiple indoor environments using instruction-conditioned waypoint navigation.