发表机构
University of Hamburg; Aswan University; University of Southampton(汉堡大学; 阿斯旺大学; 南安普顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DreamFormer利用Transformer世界模型在潜在想象中模仿专家演示,实现语言条件机器人操作,在CALVIN基准上超越现有方法,并显著提升零样本迁移性能。
AI 中文摘要
我们提出了DreamFormer,一种基于模型的智能体,通过学习到的世界模型的潜在想象中模仿专家演示,来获取语言条件下的多任务技能。DreamFormer首先从非结构化玩耍数据中学习一个任务无关的Transformer世界模型,然后通过优化内在奖励来获取特定任务的行为,该奖励使智能体生成的轨迹与潜在空间中的专家演示对齐。由于策略在想象内部进行在线训练,它在训练过程中会暴露于自身的错误,从而减轻了离线行为克隆固有的协变量偏移。为了使长时程想象变得可行,DreamFormer将高分辨率多视角机器人观测编码为单个输入令牌,避免了空间下采样以及先前Transformer世界模型使用的多令牌表示。在长时程CALVIN基准上,DreamFormer在单环境评估中优于可比较的基于模型的智能体LUMOS(每个五链平均完成2.52个任务对比2.34个),同时保持想象滚动可行。与行为克隆基线HULC相比,在未见环境的零样本迁移上,其性能几乎翻倍(1.30对比0.67),表明世界模型学习的动态比直接克隆的策略更容易迁移。这与生物智能体中内部模型的作用一致,其中环境动态模型支持在未先前遇到的情况下的行为。
英文摘要
We introduce DreamFormer, a model-based agent that acquires language-conditioned, multi-task skills by imitating expert demonstrations within the latent imagination of a learned world model. DreamFormer first learns a task-agnostic Transformer world model from unstructured play data, then acquires task-specific behaviors by optimizing an intrinsic reward that aligns agent-generated rollouts with expert demonstrations in latent space. Since the policy is trained on-policy inside imagination, it is exposed to its own errors during training, mitigating the covariate shift inherent to offline behavioral cloning. To make long-horizon imagination affordable, DreamFormer encodes a high-resolution multi-view robotic observation into a single input token, avoiding both spatial downsampling and the multi-token representations used by prior Transformer world models. On the long-horizon CALVIN benchmark, DreamFormer outperforms LUMOS, the comparable model-based agent, on single-environment evaluation (2.52 vs 2.34 average tasks completed per chain of five) while keeping imagination rollouts tractable. Against HULC, the behavior cloning baseline, it nearly doubles performance on zero-shot transfer to an unseen environment (1.30 vs 0.67), indicating that dynamics learned by the world model transfer more readily than a directly cloned policy. This is consistent with the role attributed to internal models in biological agents, where a model of environment dynamics supports behavior in situations not previously encountered.
Commentssubmitted to IEEE Transactions on Cognitive and Developmental Systems (TCDS)