Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
机构 * Shanghai Jiao Tong University(上海交通大学) ; UC Berkeley(伯克利大学) ; UNC, Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Shanghai Jiao Tong University(上海交通大学) ; UC Berkeley(伯克利大学) ; UNC, Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract)
机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) ; University of Surrey(Surrey大学) ; School of Computer Science, Wuhan University(武汉大学计算机学院) ; School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) ; College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) ; Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables
机构 * Sangho Lee 1,2(Sangho Lee 教授) ; Il Yong Chun 1,3(Il Yong Chun 教授) ; Hogun Park 1(Hogun Park 教授)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted to the AAAI 2025 Main Technical Track. This is an extended version of the original submission
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Purdue University(普渡大学) ; University at Albany(阿尔巴尼大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL