VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
专题命中 视觉推理 :VLM(title,abstract);vision language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :VLM(title,abstract);vision language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 视觉推理 :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments Code and data are available at https://zoezheng126.github.io/STLLM-website/
机构 * The University of Hong Kong(香港大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
机构 * X-LANCE Lab, School of Computer Science, Key Laboratory of Artificial Intelligence\ of Education, Shanghai Jiao Tong University, Shanghai 200240 , China ; Jiangsu Key Lab of Language Computing, Suzhou 215123 , China ; College of Computing ; Data Science, Nanyang Technological University, Singapore 639798 , Singapore ; Suzhou Laboratory, Suzhou 215123 , China
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 24 pages, 19 figures, 10 tables. Details and access are available at: https://OpenDFM.github.io/MULTI-Benchmark/
Journal ref Sci. China Inf. Sci. 68, 200107 (2025)
机构 * IIIS, Tsinghua University(清华大学智能技术学部)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; AgiBot ; School of Computer Science, Peking University(北京大学计算机学院)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract)
机构 * Tsinghua University(清华大学) ; ByteDance Seed(字节跳动种子) ; Princeton University(普林斯顿大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI
Comments I would like to formally request the withdrawal of my manuscript from arXiv. After a further internal review, I realized that the dataset used in this study contains personal or sensitive information that may inadvertently compromise individuals' privacy
机构 * University of California, Merced(加州大学默塞德分校) ; University of Queensland(昆士兰大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI