InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.AI
Comments 14 pages, 9 figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :vision-language model(abstract);grounding(abstract);分类 cs.AI
Comments 14 pages, 9 figures
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted by ICCV2025
机构 * New York University(纽约大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * 1 Taobao \& Tmall Group of Alibaba, 2 Institute of Software, Chinese Academy of Science, 3 University of Chinese Academy of Sciences, 4 Renmin University of China, 5 Informatics Department, PUC-Rio
专题命中 视觉推理 :vision language model(abstract);visual reasoning(abstract);分类 cs.AI
Comments 48 pages
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构 * Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025. arXiv admin note: text overlap with arXiv:2311.06602 by other authors
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; Xiamen Unisound Intelligence Technology Co., Ltd(厦门Unisound智能科技有限公司) ; Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室) ; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),中华人民共和国文化和旅游部,中国)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV
Comments Accepted by IJCAI 2025
机构 * Sun Yat-sen University(中山大学) ; Peng Cheng Laboratory(鹏城实验室) ; OPPO AI Center(OPPO人工智能中心) ; Research Institute(研究 institute) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Harbin Institute of Technology(哈尔滨工业大学) ; Shenzhen Key Laboratory of Digital Living Network and Content Service(深圳数字生活网络与内容服务重点实验室) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments published in ICCV 2025
机构 * Artificial Intelligence and Learning Systems Laboratory, National Technical University of Athens(人工智能与学习系统实验室,国家技术大学)
专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV
机构 * Singapore Institute of Management(新加坡管理学院)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.LG
机构 * Department of Computer Science University of Bari Aldo Moro(计算机科学系巴里大学Aldo Moro)
专题命中 视觉推理 :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * ShanghaiTech University(上海科技大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.LG
Comments Accepted by ICCV2025
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) ; MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) ; DAMO Academy, Alibaba Group, Hangzhou, China(阿里云达摩院) ; Hupan Lab, Hangzhou, China(杭州华普实验室)
专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * Soochow University(苏州大学) ; Microsoft(微软) ; Fudan University(复旦大学) ; Shanghai Jiao Tong University(上海交通大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Sun Yat-sen University(中山大学) ; Huazhong University of Science and Technology(华中科技大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV
Comments Work in progress
机构 * Wuhan University(武汉大学) ; Bytedance Seed(字节跳动种子) ; Peking University(北京大学) ; Zhejiang University(浙江大学) ; SJTU(上海交通大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ICCV2025
机构 * vivo AI Lab(vivo人工智能实验室)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构 * UC Berkeley(伯克利大学) ; HKU(香港大学) ; Adobe(Adobe公司)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Project page: https://danielchyeh.github.io/x-planner/
机构 * WeChat AI, Tencent Inc, China(腾讯公司)
专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV
机构 * School of Electrical and Computer Engineering, Peking University(北京大学电子工程学院) ; Hupan Lab(鸿篇实验室) ; DAMO Academy, Alibaba Group(阿里云达摩院) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * HKUST(香港科技大学) ; Dartmouth College(达特茅斯学院)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Project page: https://spatctxvlm.github.io/project_page/
机构 * TikTok, Inc.(字节跳动公司)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
Comments 10 pages, 5 figures
机构 * Department of Mechanical and Mechatronics Engineering, Faculty of Engineering and Design, University of Auckland(机械与机电工程系,工程与设计学院,奥克兰大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
Comments Submitted to JMS(March 2025)
机构 * Xiamen University(厦门大学) ; Tencent(腾讯) ; Anyang Normal University(安阳师范学院)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to ICCV 2025
机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州))
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Technical Report
机构 * University of Science and Technology of China(中国科学技术大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * FNii-Shenzhen, CUHKSZ(FNii深圳,香港科技大学)
专题命中 视觉推理 :visual question answering(abstract);grounding(abstract);分类 cs.CV
机构 * New York University Shanghai(纽约大学上海分校) ; New York University(纽约大学) ; Honda Research(本田研究)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * Monash University(墨尔本大学) ; Curtin University(Curtin大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; ByteDance Seed(字节跳动种子) ; The Chinese University of Hong Kong(香港中文大学) ; City University of Hong Kong(香港城市大学) ; BIAI-ZJUT ; Zhejiang University of Technology(浙江工业大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV