OREN-X:用于实时多模态建图的八叉树残差网络
OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping
- University of California San Diego(加州大学圣迭戈分校)
- VinMotion
- Parsons(帕森斯公司)
- U.S. DEVCOM Army Research Laboratory(美国DEVCOM陆军研究实验室)
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OREN-X 提出基于八叉树的共享数据结构,统一存储几何、辐射和视觉-语言特征,实现实时多模态建图,并通过跨模态协同和字典学习提升精度与压缩率。
AI中文摘要:
为了实现长时程的通用自主性,机器人需要维护支持多种任务的空间环境信息:用于规划和控制的几何信息、用于渲染和重定位的辐射信息,以及用于开放词汇语义 grounding 的视觉-语言特征。现有方法分别表示和估计每种模态,从而成倍增加内存和计算成本,同时放弃了表示之间潜在的协同效应。我们开发了 OREN-X,一种在线建图方法,使用三维空间中的八叉树作为共享数据结构来索引和存储多模态场,捕获几何、辐射和视觉-语言信息。OREN-X 提供了这些数据在显式/隐式和完整/压缩形式下的高效统一存储和检索。我们的统一表示产生了跨模态协同效应:SDF 估计通过占用率和辐射得到锐化,而基于 GPU 的射线-八叉树遍历和八叉树查询实现了实时渲染。我们还使用在线字典学习来压缩视觉-语言特征,将其缩小到全逐顶点存储的 3.7 倍以下,同时提高了查询准确率。在 Replica 数据集上,OREN-X 实时建图(SDF 达到 80+ fps,所有四种模态达到 30+ fps),将近表面 SDF 准确率比单模态基线提高了 33%,并将平均开放词汇 3D mIoU 比最佳先前方法提高了 71%,平均准确率提高了 61%。
英文摘要:
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.