A Matter of Time: Revealing the Structure of Time in Vision-Language Models
机构 * St. Pölten University of Applied Sciences(施普伦特应用科学大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * St. Pölten University of Applied Sciences(施普伦特应用科学大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
机构 * Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能研究中心、剑桥大学) ; Department of Engineering, University of Cambridge(工程系、剑桥大学) ; Department of Psychology, University of Cambridge(心理学系、剑桥大学) ; Department of Computer Science, University of Cambridge(计算机科学系、剑桥大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Preprint
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) ; Institute of Artificial Intelligence (TeleAI), China Telecom, P. R. China(人工智能研究所(TeleAI),中国电信,中华人民共和国) ; College of Computer Science, Wuhan University(计算机科学学院,武汉大学) ; School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院,光学与电子学(iOPEN),西北工业大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV