ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
专题命中 幻觉与鲁棒性 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 幻觉与鲁棒性 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
专题命中 幻觉与鲁棒性 :vision-language model(abstract);VLM(abstract);grounding(abstract);multimodal large language model(abstract)
Comments Accepted to ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV
机构 * School of Software Technology Zhejiang University Hangzhou Zhejiang China(软件技术学院浙江大学杭州浙江中国) ; Zhejiang University(浙江大学)
专题命中 幻觉与鲁棒性 :MLLM(abstract)
Comments Accepted at ACM MM 2025 Main Conference
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) ; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) ; Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)
专题命中 VLM训练与架构 :LLaVA(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by ACL 2025 Findings
专题命中 VLM训练与架构 :LLaVA(title);vision-language model(abstract)
Comments Project page: https://ali-vosoughi.github.io/SoundCLIP/
机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) ; China Mobile Research Institute(中国移动研究院) ; Shanghai AI Lab(上海AI实验室)
专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025; Code released at https://github.com/MCG-NJU/p-MoD
机构 * LMU Munich(慕尼黑大学) ; Technical University of Munich(慕尼黑技术大学) ; Siemens AG(西门子股份公司) ; University of Science and Technology of China(中国科学技术大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ; Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院) ; University of Oxford(牛津大学)
专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments Accepted to COLM 2025
机构 * Meituan(美团) ; Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV
Comments ICCV 2025
机构 * Sapienza University of Rome(罗马大学)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.LG
Comments Accepted as Full Research Papers at CIKM 2025
机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) ; Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 8 pages, 4 figures
机构 * University of Science and Technology of China(科学技术大学) ; Fudan University(复旦大学)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与产业智能工程研究中心) ; Deep NeuroCognition Lab, I2R and CFAR, Agency for Science, Technology and Research, Singapore(深度神经认知实验室,I2R和CFAR,科技研究局,新加坡)
专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV
Comments 10 pages, 7figures
机构 * New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(模式识别新实验室,自动化研究所,中国科学院)
专题命中 其他VLM :visual language model(abstract);分类 cs.CV
Comments accepted by ICPR2024