Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) ; Wuhan AI Research, Wuhan, China(武汉人工智能研究所)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Tech report