PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 11 tables and figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 11 tables and figures
机构 * State Key Lab. LIESMARS, Wuhan University(武汉大学遥感信息与地学实验室国家重点实验室)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025 Track on Datasets and Benchmarks
机构 * Kalinga Institute of Industrial Technology (KIIT)(喀里亚理工学院) ; Indian Institute of Technology (IIT), Bhubaneswar(印度理工学院(班加罗尔))
专题命中 视觉推理 :vision-language model(title,abstract)
Comments Accepted to IJCNLP-AACL Findings 2025
机构 * Örebro University(奥雷布罗大学) ; Meta
专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.CV
Comments Accepted at ICCV 2023
机构 * Hong Kong Baptist University(香港 Baptist 大学) ; Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) ; National University of Singapore(新加坡国立大学) ; Beijing Normal University(北京师范大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI
Comments 28 pages, 14 figures, 19 tables
机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) ; DEVCOM Army Research Laboratory(国防部陆军研究实验室)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
专题命中 视觉推理 :multimodal large language model(abstract)
Journal ref C. Li, Q. Yan, M. Kim, Z. Li, Y. Xu and L. -F. Yu, "Crafting Dynamic Virtual Activities with Advanced Multimodal Models," 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 120-130