arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22175cs.LGcs.AIcs.CV

对比世界模型

Contrastive World Models

Bonnie Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对像素重建世界模型在视觉复杂环境中易受干扰的问题,提出对比世界模型,用互信息下界替代重建损失,在干扰背景下显著优于Dreamer,且训练更高效。

中文摘要 AI 辅助

通过像素重建训练的世界模型在视觉复杂环境中可能表现不佳,因为不相关信息主导了目标函数,并将模型从与规划和决策相关的信息中分散注意力。我们提出了对比世界模型(Contrastive World Models),一种无需像素重建即可学习潜在动力学模型的方法。基于Dreamer,我们将标准世界模型目标中的观测重建替换为类似Deep InfoMax的下界,该下界最大化状态-动作序列与未来观测的局部块特征之间的互信息,鼓励状态表示保留对未来具有预测性的信息,而无需模型重建视觉上不相关的细节。我们在三个视觉复杂度递增的设置中进行了小规模实验评估。我们的方法在默认设置下与Dreamer和动量预测基线相匹配,并且在引入干扰物或自然视频背景时显著优于两者,同时通过完全移除像素解码器实现了更高效的训练。我们的方法具有通用性,仅需访问状态-动作序列和未来观测,几乎不做额外假设。这些结果表明,基于对比和互信息最大化的目标是构建对视觉干扰因素鲁棒的世界模型的原则性且有前景的方向,这一特性对于将基于模型的强化学习智能体迁移到现实世界尤为重要。

英文摘要

World models trained via pixel reconstruction can struggle in visually complex environments, where irrelevant information dominates the objective and distract the model from information relevant to planning and control. We present Contrastive World Models, an approach for learning latent dynamics models without pixel reconstruction. Building on Dreamer, we replace observation reconstruction in the standard world model objective with a Deep InfoMax-like lower bound that maximizes the mutual information between state-action sequences and local patch features of future observations, encouraging state representations to retain information that is predictive of the future without requiring the model to reconstruct visually irrelevant details. We evaluate our approach in small-scale experiments across three settings of increasing visual complexity. Our method matches Dreamer and a momentum prediction baseline in the default setting, and substantially outperforms both once distractors or natural video backgrounds are introduced, while also training more efficiently by removing the pixel decoder entirely. Our approach is general and makes minimal assumptions beyond access to state-action sequences and future observations. These results suggest that contrastive, infomax-based objectives are a principled and promising direction for building world models that are robust to visual nuisance factors, a property particularly relevant for transferring model-based RL agents to the real world.

↑