发表机构
University of Hamburg(汉堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VLA-Dreamer架构,在VLA视觉编码器嵌入空间训练世界模型,以提升样本效率并实现短期规划,弥补VLA缺乏显式世界模型的局限。
AI 中文摘要
视觉-语言-动作模型(VLAs)虽然在机器人控制方面展现出强大潜力,但需要大量高质量的模仿学习数据。此外,缺乏显式世界模型进一步对其控制能力产生质疑。在这篇概念论文中,我们提出了一种新颖架构,通过在VLA视觉编码器的嵌入空间上训练预测性世界模型来解决VLA中的样本效率问题。我们假设这些嵌入与动作相关且可用于未来预测。为此,我们建议使用所提出的架构来研究这些嵌入基于动作预测未来的能力,因为无法做到这一点将标志着VLA架构的一个关键局限性:缺乏无损隐式世界模型来模拟真实世界动态。所提出的架构不同于标准世界模型动态,因为损失来自嵌入空间而非像素空间,类似于联合嵌入预测架构。此外,训练好的世界模型可用于短期规划任务,通过给定目标图像采样VLA动作。我们旨在检查VLA中视觉嵌入的丰富性,并通过一个在推理时也能生成计划的世界模型来降低其高数据需求。
英文摘要
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.