arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越单一潜在空间:用于长时程规划的双潜在世界模型

Beyond a single latent space: a dual-latent world model for long-horizon planning

Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang

arXiv 2609.37644首次发表:更新:

发表机构

Nanjing University; Shenzhen University of Advanced Technology; Shanghai Jiao Tong University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(南京大学; 深圳先进技术大学; 上海交通大学; 中国科学院深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在世界模型在长时程规划中的误差累积与目标判别问题,提出双潜在世界模型(Dual-WM)及加权展开表示学习(LoRe),分离局部执行与长程规划,在五个视觉控制任务上显著提升成功率。

AI 中文摘要

潜在世界模型尽管短期预测准确,但在长时程规划中常常遇到困难。递归展开会累积误差,而高维潜在空间中的距离集中会削弱目标判别能力。我们提出了双潜在世界模型(Dual-WM),通过不同的状态表示和动力学模型将局部执行与长程规划分离。低层模型预测动作条件转移,而高层模型使用学习到的宏动作在更长的时间跨度上进行规划。我们还提出了带加权展开的长时程表示学习(LoRe),该方法在两个层级上监督自生成预测。对递归误差传播的分析促使我们采用指数时域权重,并为两个时间尺度设置不同的衰减率。在规划过程中,高层模型生成潜在子目标,低层模型将其细化为动作以精确执行。我们在五个目标条件视觉控制任务上从头评估了Dual-WM,并与无actor引导提案的任务最强基线进行了比较。在目标偏移为50和100环境步时,平均成功率分别从75.9%提高到84.4%,从61.4%提高到69.5%。在偏移100时,Dual-WM在所有五个任务上均优于这些基线,并将平均成功率比LeWM提高了30.8个百分点。消融实验和支持性分析提供了证据,表明在目标评估中表示更具信息性,在递归预测下一致性更强。这些结果凸显了分离时间角色和在多个时域上进行训练对于可靠的潜在规划的价值。我们的核心实现可在以下https URL获取。

英文摘要

Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.

Comments31 pages, 22 figures, 9 tables. Main text: 9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑