arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21502cs.CVcs.RO

自适应世界记忆3D基础模型,用于可扩展的3D建图、定位与渲染

Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering

  • Shanghai Jiao Tong University(上海交通大学)
  • State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis(航空电子综合与航空系统体系综合国家重点实验室)
  • Shanghai Key Laboratory of Navigation and Location Based Services(上海市导航与位置服务重点实验室)
  • Nanyang Technological University(南洋理工大学)
  • University of Technology Nuremberg(纽伦堡工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianchen Deng, Guole Shen, Yilin Shen, Wenhua Wu, Yilin Fang, Ziqi Ma, Tianjun Zhang, Shenghai Yuan, Wolfram Burgard, Hesheng Wang

AI总结:

提出一种以自适应世界记忆为核心的3D基础模型,通过门控更新与时空调控实现可扩展建图、定位与高斯渲染,在轨迹精度、重建完整性和渲染质量上超越现有基线。

AI中文摘要:

现有的3D基础模型能够从RGB图像进行泛化的几何推理,但在持久记忆、可扩展性以及可渲染场景建模方面仍存在局限。我们提出了一种以记忆为中心的3D基础模型,用于可扩展的机器人定位、重建和高斯渲染。其核心是一种自适应世界记忆机制,该机制结合了基于Transformer的门控更新与测试时的时间-空间调控。学习到的门控控制循环记忆传播,而时间状态演化与空间观测-状态一致性则对长图像序列上的逐令牌更新与遗忘进行调控。为支持大规模建图,我们将记忆组织为局部子地图,并集成渐进式建图与跟踪、回环检测以及基于SL(4)的全局优化,以保持局部精度与全局一致性。一个高斯重建头将记忆增强的特征解码为可渲染的图元,从而在单一模型中统一了相机位姿估计、密集点云重建和照片级真实感渲染。在公开基准和来自多种机器人平台的自采集数据集上的实验表明,与现有的3D基础重建和SLAM基线相比,我们的方法在轨迹精度、重建完整性和渲染质量方面均有提升。这些结果支持自适应记忆作为持久机器人世界建模的基础。数据集和代码将在 \n\href{ this https URL }{ this https URL } 公开提供。

英文摘要:

Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global consistency. A Gaussian reconstruction head decodes memory-enhanced features into renderable primitives, unifying camera pose estimation, dense point-cloud reconstruction, and photorealistic rendering within a single model. Experiments on public benchmarks and self-collected datasets from diverse robotic platforms demonstrate improved trajectory accuracy, reconstruction completeness, and rendering quality over existing 3D foundation reconstruction and SLAM baselines. These results support adaptive memory as a foundation for persistent robotic world modeling. The dataset and code will be made publicly available at \href{https://github.com/dtc111111/AWM-3DFM}{https://github.com/dtc111111/AWM-3DFM}.

↑