From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
从像素到谓词:通过预训练视觉-语言模型学习符号世界模型
机构 * MIT(麻省理工学院) ; Princeton University(普林斯顿大学) ; University of Cambridge(剑桥大学) ; RAI Institute(RAI研究院)
AI总结 通过预训练视觉-语言模型学习符号世界模型,以实现复杂机器人领域中长周期决策制定的零样本泛化。
Comments A version of this paper appears in the official proceedings of RA-L, Volume 11, Issue 4