arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NaviScale:生成用于目标导航的大规模语义地图数据集

NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation

Chuanlin Lan, Yanwei Zheng, Yuxi Jing, Weijian Liu, Zhitong Zhou, Jiarui Fan, Fuzhen Zhuang, Xiao Zhang, Dongxiao Yu

arXiv 2609.27218首次发表:更新:

发表机构

Shandong University; Beihang University(山东大学; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NaviScale通过组合真实住宅平面图与房间级语义地图,生成大规模训练数据,提升目标导航性能,在HM3D和MP3D上分别达到64.3%和43.1%的成功率。

AI 中文摘要

具身导航需要能够泛化到未见环境的空间表示,然而从真实3D环境中收集大量带标注数据十分困难。我们提出NaviScale,用于基于语义地图的目标导航(ObjectNav),其预测器可以在部分和完整语义地图对上训练,而无需为每个训练样本重建完整的3D环境。该框架通过将真实住宅的平面图与从MP3D和HM3DSem中提取的房间级语义和障碍物地图组合,生成大规模语义地图训练数据。NaviScale通过两种方式增加数据多样性:房间间缩放增加了平面图级别的结构多样性,而房间内缩放则用按房间类别匹配的不同房间地图组合填充每个固定平面图。通过射线投射的可见性(VisRC)将组合地图转换为考虑视野、感知范围和遮挡的部分观测。生成的数据集包含从24,000个平面图(关联12,794个属性)生成的192,000个语义地图。使用300k次训练迭代以及本文所述的训练和推理设置,系统在HM3D上达到64.3%的SR和34.8%的SPL,在MP3D上达到43.1%的SR和16.8%的SPL,且无需更改预测架构。额外的实验评估了组合地图的质量、语义分割错误的影响以及在物理机器人上的部署。

英文摘要

Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.

Comments14 pages, 8 figures; includes supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑