Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
Goal-VLA: 图像生成视觉语言模型作为对象中心世界模型,赋能零样本机器人操作
Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
The HKU Musketeers Foundation Institute of Data Science, The University of Hong Kong(香港大学数据科学研究院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
机构
*
University of Washington(华盛顿大学)
;
National University of Singapore(新加坡国立大学)
;
Clemson University(克莱姆森大学)
;
Drexel University(德雷塞尔大学)
;
Microsoft Research(微软研究院)