CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室) ; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室) ; IHPC, Agency for Science, Technology and Research, Singapore(新加坡科技研究局IHPC) ; Show Lab, National University of Singapore(新加坡国立大学Show实验室) ; College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院)
专题命中 视觉推理 :visual language model(title);vision language model(abstract);LLaVA(abstract);InternVL(abstract)