11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
Chengzu Li, Wenshan Wu, Huanyu Zhang, Qingtao Li, Zeyu Gao, Yan Xia, José Hernández-Orallo, Ivan Vulić, Furu Wei
机构
*
Microsoft Research(微软研究院)
;
Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)
;
Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
;
Department of Oncology, University of Cambridge(癌症部门,剑桥大学)
;
Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能中心,剑桥大学)
;
VRAIN, Universitat Politècnica de València(VRAIN,巴塞罗那理工大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments9 pages, 4 figures (22 pages, 7 figures, 7 tables including references and appendices)
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah
机构
*
Computer Engineering Department, Sharif University of Technology, Iran(谢尔盖大学计算机工程系,伊朗)
;
Qatar Computing Research Institute, Qatar(卡塔尔计算研究所,卡塔尔)
;
Computer Engineering Department, University of Isfahan, Iran(伊斯法罕大学计算机工程系,伊朗)
;
Independent Researcher(独立研究者)
机构
*
Pratt School of Engineering(普拉特工程学院)
;
Duke University(杜克大学)
;
School of Computer Science(计算机科学学院)
;
Northeast Electric Power University(东北电力大学)
;
Department of Information and Communication Engineering(信息与通信工程系)
;
Tongji University(同济大学)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
School of International Education(国际教育学院)
;
Beijing University of Chemical Technology(北京化工大学)
;
East China Normal University(华东师范大学)
;
Information Hub(信息枢纽)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Chinese Academy of Science(中国科学院)
专题命中
视觉推理
:vision language model(abstract);VLM(abstract);分类 cs.AI、cs.LG
Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
Antonia Wüst, Tim Woydt, Lukas Helff, Inga Ibs, Wolfgang Stammer, Devendra S. Dhami, Constantin A. Rothkopf, Kristian Kersting
机构
*
AIML Lab, TU Darmstadt(AIML实验室,德累斯顿技术大学)
;
Hessian Center for AI (hessian.ai)(黑森人工智能中心)
;
Institute of Psychology, TU Darmstadt(心理学研究所,德累斯顿技术大学)
;
Centre for Cognitive Science, TU Darmstadt(认知科学中心,德累斯顿技术大学)
;
German Center for AI (DFKI)(德国人工智能中心)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
Huanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, Yuxin Song, Wenhao Wu, Dacheng Tao
机构
*
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
Baidu Inc.(百度公司)
;
University of California(加州大学)
;
University of Science and Technology of China(中国科学技术大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
Alexander Gambashidze, Konstantin Sobolev, Andrey Kuznetsov, Ivan Oseledets
专题命中
视觉推理
:VLM(abstract);visual language model(abstract);分类 cs.CV、cs.AI
CommentsWe are withdrawing this paper because the main contributions and methodology have significantly changed after further research and experimental updates. The current version no longer reflects our results and main contribution / topic
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Yufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue, Ruipu Luo, Zhenghao Chen, Can Zhang, Yifan Li, Zhentao He, Zheming Yang, Ming Tang, Minghui Qiu, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院)
;
ByteDance(字节跳动)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Wuhan AI Research(武汉人工智能研究所)
;
Xi’an Jiaotong University(西安交通大学)
;
Renmin University of China(中国人民大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
Zhuobai Dong, Junchao Yi, Ziyuan Zheng, Haochen Han, Xiangxi Zheng, Alex Jinpeng Wang, Fangming Liu, Linjie Li
机构
*
Central South University(中南大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Nanjing University(南京大学)
;
Microsoft(微软公司)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI