What is the Visual Cognition Gap between Humans and Multimodal LLMs?
机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) ; College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) ; Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) ; Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) ; HKUST (Guangzhou)(香港科技大学(广州)) ; Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)
专题命中 视觉定位与Grounding :vision language model(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments COLM 2025