Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
Antonia Wüst, Tim Woydt, Lukas Helff, Inga Ibs, Wolfgang Stammer, Devendra S. Dhami, Constantin A. Rothkopf, Kristian Kersting
机构
*
AIML Lab, TU Darmstadt(AIML实验室,德累斯顿技术大学)
;
Hessian Center for AI (hessian.ai)(黑森人工智能中心)
;
Institute of Psychology, TU Darmstadt(心理学研究所,德累斯顿技术大学)
;
Centre for Cognitive Science, TU Darmstadt(认知科学中心,德累斯顿技术大学)
;
German Center for AI (DFKI)(德国人工智能中心)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
Huanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, Yuxin Song, Wenhao Wu, Dacheng Tao
机构
*
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
Baidu Inc.(百度公司)
;
University of California(加州大学)
;
University of Science and Technology of China(中国科学技术大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
Alexander Gambashidze, Konstantin Sobolev, Andrey Kuznetsov, Ivan Oseledets
专题命中
视觉推理
:VLM(abstract);visual language model(abstract);分类 cs.CV、cs.AI
CommentsWe are withdrawing this paper because the main contributions and methodology have significantly changed after further research and experimental updates. The current version no longer reflects our results and main contribution / topic
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Yufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue, Ruipu Luo, Zhenghao Chen, Can Zhang, Yifan Li, Zhentao He, Zheming Yang, Ming Tang, Minghui Qiu, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,自动化研究所,中国科学院)
;
ByteDance(字节跳动)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Wuhan AI Research(武汉人工智能研究所)
;
Xi’an Jiaotong University(西安交通大学)
;
Renmin University of China(中国人民大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
Zhuobai Dong, Junchao Yi, Ziyuan Zheng, Haochen Han, Xiangxi Zheng, Alex Jinpeng Wang, Fangming Liu, Linjie Li
机构
*
Central South University(中南大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Nanjing University(南京大学)
;
Microsoft(微软公司)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI