Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China ; School of Cyber Science ; Engineering, Nanjing University of Science ; VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China ; Key Laboratory of Ethnic Language Intelligent Analysis ; Security Governance of MOE, Minzu University of China Beijing China ; Institute of Information Engineering, Chinese Academy of Sciences ; School of Cyber Security, University of Chinese Academy of Sciences ; VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University ; Security Governance of MOE, Minzu University of China
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
Comments Accepted by 2025 ACM MM
专题命中 视觉问答 :visual question answering(abstract)
Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)