Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
Goal-VLA: 图像生成视觉语言模型作为对象中心世界模型,赋能零样本机器人操作
Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
The HKU Musketeers Foundation Institute of Data Science, The University of Hong Kong(香港大学数据科学研究院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
机构
*
University of Washington(华盛顿大学)
;
National University of Singapore(新加坡国立大学)
;
Clemson University(克莱姆森大学)
;
Drexel University(德雷塞尔大学)
;
Microsoft Research(微软研究院)
Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
梦想回溯:基于想象的体验检索用于记忆持久的视觉与语言导航
Yunzhe Xu, Yiyuan Pan, Zhe Liu
机构
*
National Key Laboratory of Human Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(西安交通大学人机混合增强智能全国重点实验室)
;
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
机构
*
Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建筑工程学院)
;
Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系)
;
Google(谷歌)