arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CausalDreamer:学习具有潜在解耦的预测世界模型

CausalDreamer: Learning Predictive World Models with Latent Disentanglement

Prince Jha, Nils Lukas, Kun Zhang, Salem Lahlou

arXiv 2610.12016首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); Carnegie Mellon University(穆罕默德·本·扎耶德人工智能大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CausalDreamer通过冻结Dreamer 4的分词器并将其潜在向量重新编码为可控性和奖励相关性轴的分解表示,微调动力学模型,在MMBench2任务上提升了规划性能,能区分奖励无关与相关变化。

AI 中文摘要

用于控制的世界模型必须捕捉环境中哪些方面会响应智能体的动作,哪些与奖励相关。生成式世界模型如Dreamer 4由视频分词器和动力学模型组成,视频分词器将每一帧编码为潜在向量,动力学模型经预训练后可根据过去的潜在向量和动作预测未来的潜在向量。然而,该分词器是通过重构目标进行训练的,没有动作或奖励监督,因此其潜在向量无法提供明确的机制来区分可控、不可控、与奖励相关及与奖励无关的信息。我们提出CausalDreamer,该方法保持分词器冻结,将其潜在向量重新编码为沿两个轴分解的四组表示:可控性轴(仅两个可控组接收动作输入)和奖励相关性轴(通过从两个与奖励相关的组预测奖励来学习),随后对预训练的动力学模型进行微调以预测该分解表示。我们在20个MMBench2任务上通过模型预测规划评估了CausalDreamer及其起始的预训练世界模型:10个训练期间见过的干净任务,10个未见过的任务,其中6个是干净任务的变体(背景、对象或迷宫布局改变),4个是新环境。我们对回报进行归一化,使得采取均匀随机动作的策略得分为0,专家策略得分为1。在干净任务上,CausalDreamer的归一化得分比预训练世界模型高14%(0.199 vs. 0.175);在变体任务上,得分高25%(0.307 vs. 0.246);而在新环境中,两个模型的得分均未显著高于随机策略。此外,我们的分析表明,该分解表示可将与奖励无关的变化(如背景改变)与与奖励相关的组区分开。

英文摘要

World models for control must capture which aspects of the environment respond to the agent's actions and which are relevant to reward. Generative world models such as Dreamer 4 consist of a video tokenizer, which encodes each frame into a latent, and a dynamics model, which is pretrained to predict future latents from past latents and actions. Yet the tokenizer is trained with a reconstruction objective, without action or reward supervision, so its latent provides no explicit mechanism to separate controllable, uncontrollable, reward-relevant, and reward-irrelevant information. We propose \textit{CausalDreamer}, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups. The pretrained dynamics model is then fine-tuned to predict the factored representation. We evaluate \textit{CausalDreamer} and the pretrained world model it starts from with model-predictive planning on 20 MMBench2 tasks: 10 clean tasks seen during training and 10 unseen tasks, of which 6 are manipulated variants of clean tasks with a changed background, object, or maze layout, and 4 are new environments. We normalize returns so that a policy taking uniformly random actions scores 0 and an expert scores 1. \textit{CausalDreamer} achieves a 14\% higher normalized score than the pretrained world model on the clean tasks (0.199 vs.\ 0.175) and a 25\% higher score on the manipulated variants (0.307 vs.\ 0.246), while neither model scores meaningfully above the random policy in the new environments. Additionally, our analysis shows that the factored representation separates reward-irrelevant changes, such as a changed background, from its reward-relevant groups.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑