arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HIGS:用于几何与语义场景联合理解的分层隐式网格

HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding

Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov

arXiv 2609.38620首次发表:更新:

发表机构

University of California San Diego; Microsoft Spatial AI Lab; University of Michigan; Southern University of Science and Technology; Norwegian University of Science and Technology; Microsoft Research(加州大学圣地亚哥分校; 微软空间人工智能实验室; 密歇根大学; 南方科技大学; 挪威科技大学; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HIGS分层隐式网格,通过多分辨率子地图和统一查询解码机制,实现高效可扩展的几何与语义场景联合理解,显著提升计算效率并保持高精度。

AI 中文摘要

神经隐式表示通过使机器人能够构建连续、可微分且高保真的3D地图,对场景重建产生了显著影响。大多数现有工作侧重于几何重建,缺乏用于高层空间理解和任务规划的语义信息。此外,随着环境规模和复杂性的增加,神经表示在后端优化中面临保持计算效率的挑战。为解决这两个挑战,我们引入了一种分层神经场,利用多分辨率子地图实现高效且可扩展的隐式表示,并采用统一的查询和解码机制来支持几何和语义特征。具体而言,可学习的地图特征可通过查询和解码过程转换为输出,适用于训练和推理。对于大规模表示,我们将场景分解为重叠的子地图,并在每个局部子地图内进行分层优化,从而实现可扩展的计算。为进一步提高效率,我们设计了特征编码器,用于预测初始分层网格特征,从而大幅减少从零优化子地图特征所需的时间。为纠正子地图之间的估计漂移,我们完全在隐式特征空间内对其进行对齐和融合,通过避免解码最终输出而实现显著加速。基于这种高效的分层表示,我们将几何特征和视觉-语言潜在特征嵌入到地图中,并在有符号距离场(SDF)构建和开放词汇对象定位上进行了演示。我们的方法显著提高了计算和内存效率,保持了高估计精度,并在大规模真实世界基准上赋予机器人空间感知能力。

英文摘要

Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-fidelity 3D maps. Most existing works focus on geometric reconstruction and lack semantic information for high-level spatial understanding and task planning. Also, as the scale and complexity of the environment increase, neural representations face the challenge of maintaining computational efficiency in back-end optimization. To resolve these two challenges, we introduce a hierarchical neural field that leverages multiresolution submaps to achieve an efficient and scalable implicit representation, and a unified query and decoding mechanism to support both geometric and semantic features. More specifically, the learnable map features can be converted to the output with the query and decoding process for both training and inference. For large-scale representation, we decompose a scene into overlapping submaps and do hierarchical optimization within each local submap, thus enabling scalable computation. To further improve efficiency, we design feature encoders that predict initial hierarchical grid features to substantially reduce the time needed to optimize the submap features from scratch. To correct estimation drift among submaps, we align and fuse them entirely within the implicit feature space, leading to substantial acceleration by avoiding the need to decode the final output. Building upon this efficient hierarchical representation, we embed both geometric features and vision-language latent features into the map, and demonstrate it on both Signed Distance Field (SDF) construction and open-vocabulary object grounding. Our approach significantly improves computation and memory efficiency, maintains high estimation accuracy, and endows the robot with spatial awareness on large-scale real-world benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑