VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
机构 * Tsinghua University(清华大学) ; Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; ByteDance China(字节跳动中国) ; Peng Cheng Laboratory(鹏城实验室) ; Wuhan AI Research(武汉人工智能研究)
专题命中 多模态生成 :multimodal(title);分类 cs.CV