arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用潜世界模型预测后果并强化导航策略

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

Zengmao Wang, Wei Gao, Shuhan Shen

arXiv 2608.26190首次发表:更新:

发表机构

School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院; 中国科学院自动化研究所; 中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出用于机器人导航的兼容性预测潜世界模型(LWM),通过预测动作条件下的潜特征兼容性评估动作后果,可从无标注视频监督策略学习并经强化学习改进,在多机器人导航数据集上性能优于现有方法。

AI 中文摘要

世界模型使智能体能够推理未来结果并基于状态转移知识学习策略,但现有方法主要聚焦于重构未来观测或特征,这引入了不必要的复杂性并限制了其在决策中的有效性。本研究中,我们提出一种用于机器人导航的兼容性预测潜世界模型(Latent World Model,LWM),该模型预测动作条件下的潜特征兼容性而非重构观测。我们的核心见解是空间邻近性与潜特征相似性相关,使得可以直接在潜空间中评估动作后果。为支持反事实训练,我们的模型利用从轨迹中采样的动作序列,并学习预测哪些序列能更接近目标。此外,我们展示了所学的世界模型如何从无标注视频数据中监督策略学习,并完全在世界模型内通过强化学习进一步改进策略。这种想象驱动的框架消除了对动作标注和额外环境交互的需求。在多个真实世界机器人导航数据集上的大量实验表明,我们的方法在预测精度、策略学习和真实世界导航性能方面显著优于现有世界模型和模仿学习方法。代码、预训练模型及额外材料可在该https URL获取。

英文摘要

World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or features, which introduces unnecessary complexity and limits their effectiveness for decision making. In this work, we propose a compatibility prediction Latent World Model (LWM) for robot navigation that predicts action-conditioned latent feature compatibility rather than reconstructing observations. Our key insight is that spatial proximity correlates with latent feature similarity, enabling action consequences to be evaluated directly in latent space. To support counterfactual training, our model leverages action sequences sampled across trajectories and learns to predict which sequences lead closer to the goal. Furthermore, we demonstrate how the learned world model can supervise policy learning from unlabeled video data and further improve policies through reinforcement learning entirely within the world model. This imagination-driven framework eliminates the need for action annotations and additional environment interaction. Extensive experiments on multiple real-world robot navigation datasets show that our approach significantly outperforms prior world model and imitation learning methods in prediction accuracy, policy learning, and real-world navigation performance. The code, pretrained models, and additional materials are available at https://wzm206.github.io/latent-world-model-nav.

Journal refECCV 2026 (Spotlight)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑