arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09518cs.CVcs.RO

ActiveLang:基于语义不确定性引导探索的主动开放词汇三维建图

ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

  • Stevens Institute of Technology(史蒂文斯理工学院)
  • Purdue University(普渡大学)
  • Goertek Alpha Labs(歌尔阿尔法实验室)

机构由 AI 辅助整理,请以论文原文为准。

Liyan Chen, Hairong Yin, Huangying Zhan, Yi Xu, Raymond A. Yeh, Philippos Mordohai

AI总结:

ActiveLang提出一种基于语义不确定性引导探索的主动开放词汇三维建图系统,通过在线语言特征自适应和高效视点规划,在Replica和ScanNet++上显著提升分割性能并降低建图成本。

AI中文摘要:

随着机器人越来越多地协助人类完成各种任务,它们需要对其周围环境具备几何和语义两方面的理解。此外,机器人通常在陌生环境中运行,并承担新任务,而事先不知道相关概念。这促使了支持开放词汇场景理解和人机交互的语言标注三维地图的发展。我们引入了ActiveLang,一个用于主动开放词汇三维建图并带有语义不确定性引导探索的自主系统。ActiveLang在紧凑的双高斯表示上执行在线语言特征自适应,以联合重建场景几何、外观和开放词汇语义,且内存开销适中。其规划器高效地选择信息丰富的视点,从而以更少的观测和更低的计算成本实现有效建图。在Replica和ScanNet++上的实验表明,与在线和离线基线相比,在二维和三维开放词汇分割方面均有显著改进,凸显了主动探索场景能更高效地构建语言标注的三维地图。

英文摘要:

As robots increasingly assist humans with diverse tasks, they need both geometric and semantic understanding of their surroundings. Moreover, robots often operate in unfamiliar environments and take on new tasks without knowing the relevant concepts ahead of time. This motivates language-annotated 3D maps that support open-vocabulary scene understanding and human-robot interaction. We introduce ActiveLang, an autonomous system for active open-vocabulary 3D mapping with semantic-uncertainty-guided exploration. ActiveLang performs online language-feature adaptation on a compact dual-Gaussian representation to jointly reconstruct scene geometry, appearance, and open-vocabulary semantics with modest memory overhead. Its planner efficiently selects informative viewpoints, enabling effective mapping with fewer observations and lower computational cost. Experiments on Replica and ScanNet++ demonstrate substantial improvements in 2D and 3D open-vocabulary segmentation over both online and offline baselines, highlighting that actively exploring scenes builds language-annotated 3D maps more efficiently.

↑