MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments 9 pages, 6 figures, work in progress
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments 9 pages, 6 figures, work in progress
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
Comments 9 pages, 6 figures
机构 * Johns Hopkins University(约翰霍普金斯大学) ; University of North Carolina Chapel Hill(北卡罗来纳大学教堂山分校) ; University of Michigan(密歇根大学) ; University of California San Diego(加州大学圣地亚哥分校) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments Published at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment (BiAlign)
机构 * Peking University(北京大学) ; Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) ; Northwestern Polytechnical University(西北工业大学) ; Southeast University(东南大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
机构 * Johns Hopkins University(约翰霍普金斯大学) ; University of California San Diego(加州大学圣地亚哥分校) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) ; University of Michigan(密歇根大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments Published at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment (BiAlign)
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Shanghai Innovation Institute(上海创新研究院) ; USTC(中国科学技术大学) ; RIT(罗切斯特理工学院) ; HIT(哈尔滨工业大学) ; WHU(武汉大学) ; MBZUAI(马克斯·普朗克人工智能研究所) ; NUS(新加坡国立大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
Comments 35 pages, 33 figures
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments 29 pages, 16 figures
专题命中 视觉推理 :VLM(title);vision-language model(abstract);分类 cs.AI
Comments Project Web: https://robo-intention.github.io
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI
Comments 6 pages, 3 figures. Accepted, Presented and Published as part of Proceedings of the 6th International Conference on Recent Advantages in Information Technology (RAIT) 2025
Journal ref 2025 6th International Conference on Recent Advances in Information Technology (RAIT), Dhanbad, India, 2025, pp. 1-6
机构 * Xiamen University(厦门大学) ; Zhongguancun Academy(中关村学院)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments ACM MM25 has accepted this paper
机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德大学人工智能学院) ; Tianjin University(天津大学) ; Nanjing University(南京大学) ; Aalto University(艾尔沃斯大学)
专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV
机构 * Northwestern Polytechnical University(西北工业大学) ; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted By ACM MM 25
机构 * Johns Hopkins University(约翰霍普金斯大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
机构 * School of Computer Science(计算机科学学院) ; Informatics De Montfort University(信息学德蒙特福特大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Department of Pathology, Liuzhou People’s Hospital Affiliated to Guangxi Medical University(广西医科大学柳州市人民医院病理科) ; Department of Immunology, College of Basic Medical Sciences, China Medical University(中国医科大学基础医学学院免疫科) ; Greater Bay Area Center for Medical Device Evaluation and Inspection.NMPA(粤港澳大湾区医疗器械评价和检验中心.NMPA) ; Shenzhen Shengqiang Technology Co., Ltd.(深圳盛强科技有限公司) ; Department of Pathology, Chongqing University Affiliated Three Gorges Hospital(重庆大学附属第三人民医院病理科)
专题命中 视觉推理 :vision-language model(title);vision language model(abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments 15 pages, 10 figures
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
Comments COLM 2025. VisOnlyQA dataset, code, and model responses are provided at https://github.com/psunlpgroup/VisOnlyQA. Please also refer to our project website at https://visonlyqa.github.io/
机构 * York University(约克大学) ; University of Alberta(阿尔伯塔大学) ; Nanyang Technological University(南洋理工大学) ; Salesforce AI Research(Salesforce AI研究)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted at ACL 2025 Industry Track
机构 * Autolab, Westlake University(Autolab,西拉丘吉大学) ; Tianjin University(天津大学) ; Zhejiang University(浙江大学) ; Capital Normal University(首都师范大学) ; University of Hong Kong(香港大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
机构 * New York University(纽约大学) ; Yale University(耶鲁大学) ; Stanford University(斯坦福大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
Comments Project page: https://vision-x-nyu.github.io/thinking-in-space.github.io/
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI
Comments ACL 2025
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by ICCV 2025. Project page: https://zifuwan.github.io/ONLY/
机构 * HKU(香港大学) ; Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) ; CUHK(香港中文大学) ; HUST(华中科技大学)
专题命中 视觉推理 :visual reasoning(title);vision-language model(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; Alibaba Group(阿里巴巴集团)
专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Journal ref Findings of the ACL 2025
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学) ; University of Science and Technology of China(中国科学技术大学) ; Shanghai Jiao Tong University(上海交通大学) ; Donghua University(东华大学) ; Nanjing University(南京大学) ; Tsinghua University(清华大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
机构 * The Chinese University of Hong Kong(香港中文大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Tsinghua University(清华大学)
专题命中 视觉推理 :visual reasoning(title);multimodal large language model(abstract);分类 cs.CV
专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV
Comments Accepted in CVPR 2025 Workshop (CVinW)
机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(计算与智能研究院,哈尔滨工业大学,深圳) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳) ; School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, China(智能科学与工程学院,哈尔滨工业大学,深圳)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by ACL2025 main conference
机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV
Comments Accepted by Findings of ACL 2025