ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry
ChemVTS-Bench: 评估多模态大语言模型在化学中的视觉-文本-符号推理能力
Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su
机构
*
Renmin University of China(中国人民大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
South China University of Technology(华南理工大学)
;
Gaotu Techedu Inc(高图科技公司)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.AI
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
Sensetime Research(商汤科技研究院)
;
Beijing Institute of Technology(北京理工大学)
;
Shanghai AI Lab(上海人工智能实验室)
专题命中
视觉推理
:MLLM(title);multimodal large language model(abstract);分类 cs.CV
机构
*
University of Electronic Science and Technology of China(电子科学与技术大学)
;
Southwestern University of Finance and Economics(西南财经大学)
;
Tongji University(同济大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV
Comments19 pages, 11 figures. Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
Wei Zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang, Yin Zhuang, Shijian Lu, He Chen, Jun Li, Xuerui Mao
机构
*
School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
;
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学)
;
School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学)
;
State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)
$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Trishanu Das, Abhilash Nandy, Khush Bajaj, Deepiha S
机构
*
Tredence Inc.(特伦德公司)
;
Indian Institute of Technology Kharagpur(印度理工学院克拉格浦尔分校)
;
Inria Paris-Rocquencourt(巴黎-罗克琴库特研究所)
;
Rajiv Gandhi University(拉吉夫·甘地大学)
;
Tsinghua University(清华大学)
;
Palmer Research Laboratories(帕勒姆研究实验室)
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, Yuheng Li, Konstantinos Psounis, Xiaofeng Yang
机构
*
Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学)
;
Department of Computer Science and Department of Electrical and Computer Engineering, University of Southern California(计算机科学系和电气与计算机工程系,南加州大学)
;
Department of Computer Science, University of Tokyo(计算机科学系,东京大学)
;
Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学)
;
Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
Brandon James Carone, Iran R. Roman, Pablo Ripollés
机构
*
Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室)
;
Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.AI
CommentsAccepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables
机构
*
School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学)
;
State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所)
;
Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Xiamen University(厦门大学)
;
The Hong Kong University of Science and Technology(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.AI