MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
Tajamul Ashraf, Umair Nawaz, Abdelrahman M. Shaker, Rao Anwer, Philip Torr, Fahad Shahbaz Khan, Salman Khan
专题命中
视觉推理
:vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
CommentsWe have come across a recent approach that has not been properly attributed at the time of submission and compared in a fair setting. Therefore, we would like to withdraw the paper to address these concerns
MULTI: Multimodal Understanding Leaderboard with Text and Images
Zichen Zhu, Yang Xu, Lu Chen, Jingkai Yang, Yichuan Ma, Yiming Sun, Hailin Wen, Jiaqi Liu, Jinyu Cai, Yingzi Ma, Situo Zhang, Zihan Zhao, Liangtai Sun, Kai Yu
机构
*
X-LANCE Lab, School of Computer Science, Key Laboratory of Artificial Intelligence\ of Education, Shanghai Jiao Tong University, Shanghai 200240 , China
;
Jiangsu Key Lab of Language Computing, Suzhou 215123 , China
;
College of Computing
;
Data Science, Nanyang Technological University, Singapore 639798 , Singapore
;
Suzhou Laboratory, Suzhou 215123 , China
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein, Micah Goldblum
机构
*
Columbia University(哥伦比亚大学)
;
University of Maryland(马里兰大学)
;
University of Southern California(南加州大学)
;
New York University(纽约大学)
机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Nanjing University(南京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Dexmal
机构
*
SKL-MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所SKL-MAIS部门,北京)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京)
;
EACON, Fujian, China(福建EACON机构,中国)
;
School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing, China(北京科技大学自动化与电气工程学院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
Chengzu Li, Wenshan Wu, Huanyu Zhang, Qingtao Li, Zeyu Gao, Yan Xia, José Hernández-Orallo, Ivan Vulić, Furu Wei
机构
*
Microsoft Research(微软研究院)
;
Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)
;
Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
;
Department of Oncology, University of Cambridge(癌症部门,剑桥大学)
;
Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能中心,剑桥大学)
;
VRAIN, Universitat Politècnica de València(VRAIN,巴塞罗那理工大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments9 pages, 4 figures (22 pages, 7 figures, 7 tables including references and appendices)
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah
机构
*
Computer Engineering Department, Sharif University of Technology, Iran(谢尔盖大学计算机工程系,伊朗)
;
Qatar Computing Research Institute, Qatar(卡塔尔计算研究所,卡塔尔)
;
Computer Engineering Department, University of Isfahan, Iran(伊斯法罕大学计算机工程系,伊朗)
;
Independent Researcher(独立研究者)
机构
*
Pratt School of Engineering(普拉特工程学院)
;
Duke University(杜克大学)
;
School of Computer Science(计算机科学学院)
;
Northeast Electric Power University(东北电力大学)
;
Department of Information and Communication Engineering(信息与通信工程系)
;
Tongji University(同济大学)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
School of International Education(国际教育学院)
;
Beijing University of Chemical Technology(北京化工大学)
;
East China Normal University(华东师范大学)
;
Information Hub(信息枢纽)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Chinese Academy of Science(中国科学院)
专题命中
视觉推理
:vision language model(abstract);VLM(abstract);分类 cs.AI、cs.LG