KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?
KidVis: 多模态大语言模型是否具备六岁儿童的视觉感知能力?
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
AI总结 KidVis研究通过对比人类儿童与多模态大语言模型在视觉能力上的表现,揭示了当前模型在基础视觉感知上的不足。