arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14561cs.RO

GLAM:基于全局时空记忆的潜在世界模型,用于主动探索与导航

GLAM: Training a latent world model over global spatiotemporal memory for active exploration and navigation

I-Tak Ieong, Ruizhi Feng, Zhaoyang Lu, Yifei Cao, Jiayao Zhao, Leon Li, Senhua Zhu, Wenbo Ding

首次发表
浏览论文内容

中文总结 AI 辅助

提出GLAM潜在世界模型及导航系统GLAM NAV,基于全局时空记忆预测未来地图和航点,在HM3D-ObjectNav上超越BSC-Nav基线,提升主动探索与语义导航性能。

中文摘要 AI 辅助

主动探索和语义导航要求具身智能体从局部观测中构建记忆,预测观测到的空间记忆的演变如何支持未来的运动,并将该预测转化为可执行的计划。我们提出了GLAM,一种基于全局时空记忆训练的目标条件潜在世界模型,以及GLAM NAV,一个围绕该模型构建的完整导航系统。给定历史地图令牌、导航目标和当前机器人位姿,GLAM联合预测未来地图表示和以机器人为中心的航点潜在变量,从而在共享表示空间中推断未来的空间上下文和导航意图。该模型遵循类似JEPA的潜在预测范式,直接操作地图级潜在令牌而非RGB重建,并使用预训练的航点编码器-解码器在GLAM NAV中监督和解码导航计划。训练数据通过回放在HM3D v0.2场景资产上的Habitat中ObjectNav专家轨迹,并将其切片为多时间尺度预测样本而收集。在受控的HM3D-ObjectNav子集复现设置中,GLAM NAV在成功率以及按路径长度加权的成功率方面均优于复现的BSC-Nav基线。

英文摘要

Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder-decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected by replaying ObjectNav expert trajectories in Habitat over HM3D v0.2 scene assets and slicing them into multi-timescale prediction samples. On a controlled HM3D-ObjectNav subset reproduction setting, GLAM NAV improves over a reproduced BSC-Nav baseline in both success rate and success weighted by path length.

发表机构

  • Tsinghua University(清华大学)
  • Lab for Brain-Inspired Embodied Intelligence, EBKernel Technologies Co., Ltd(类脑具身智能实验室,EBKernel科技有限公司)
  • Northeastern University(东北大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑