StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Shenyang institute of computing technology, Chinese academy of sciences(沈阳计算技术研究所,中国科学院) ; Shanghai AI Laboratory(上海人工智能实验室) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
机构 * College of Computing, Georgia Institute of Technology(计算学院、佐治亚理工学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.LG
机构 * MAIS, Institute of Automation of Chinese Academy of Sciences(自动化研究所)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments 45 pages, accepted by CVPR2025
机构 * The Chinese University of Hong Kong(香港中文大学) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) ; Dexmal ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; SKL-IOTSC, CIS, University of Macau(澳门科学馆-物联网与智能系统研究中心,澳门大学)
专题命中 视觉推理 :VLM(abstract);分类 cs.CV
Comments Accepted by ICCV 2025. The code is at https://github.com/wudongming97/AffordanceNet
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Department of Electrical and Computer Engineering, Duke University(电子工程系,杜克大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments 2 pages, 4 figures
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; College of Biomedical Engineering and Instrument Science, Zhejiang University(浙江大学生物医学工程与仪器科学学院) ; Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合学院)
专题命中 视觉推理 :grounding(abstract);分类 cs.CV
机构 * Program in Neuroscience, Harvard Medical School(神经科学项目,哈佛医学院) ; Boston Children’s Hospital, Harvard Medical School(波士顿儿童医院,哈佛医学院) ; Center for Brains, Minds, and Machines(大脑、心智与机器中心) ; Edmond and Lily Safra Center for Brain Sciences, Hebrew University(埃德蒙与莉莉·萨弗脑科学中心,希伯来大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments 27 pages, 7 figures, 8 SI figures, 2 SI tables
机构 * Nanjing University(南京大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The University of Tokyo(东京大学) ; Zhejiang University(浙江大学) ; Fudan University(复旦大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * San Jose State University(圣何塞州立大学) ; Guilin University of Electronic Technology(桂林电子科技大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Mphasis Limited(默比斯有限公司)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
机构 * The University of Hong Kong(香港大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments 37 pages, 12 figures
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Microsoft(微软) ; UIUC(伊利诺伊大学香槟分校)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Nanjing University of Science and Technology(南京理工大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Code: https://dongzhang89.github.io/RGenie.github.io/
机构 * Tsinghua University(清华大学) ; Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; ByteDance China(字节跳动中国) ; Peng Cheng Laboratory(鹏城实验室) ; Wuhan AI Research(武汉人工智能研究)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * Department of Computer Science, Norwegian University of Science and Technology(计算机科学系,挪威科学技术大学) ; Department of Design, Norwegian University of Science and Technology(设计系,挪威科学技术大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI
机构 * Halıcıoğlu Data Science Institute (HDSI), University of California San Diego, La Jolla, CA, USA(Halıcıoğlu数据科学研究所(HDSI),加州大学圣地亚哥分校,拉贾拉,加州,美国)
专题命中 视觉推理 :VLM(abstract);分类 cs.AI
Comments 8 pages, ICML MAS workshop
机构 * Shanxi Province Cancer Hospital, Chinese Academy of Medical Sciences(山西省肿瘤医院、中国医学科学院) ; Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences(苏州生物医学工程与技术研究所、中国科学院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.LG
机构 * Sun Yat-sen University, Shenzhen(中山大学深圳校区) ; Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区) ; SenseTime Research(商汤科技研究院) ; Chinese University of Hong Kong(香港中文大学) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Extended Version of KptLLM. arXiv admin note: text overlap with arXiv:2411.01846
机构 * University of Science and Technology of China(中国科学技术大学) ; YuanShi Technology(元世科技) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
Comments ICCV 2025
机构 * Massachusetts Institute of Technology(麻省理工学院) ; University of Alabama(阿拉巴马大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
Comments Accepted as paper in 19th International Conference on Neurosymbolic Learning and Reasoning,NeSy 2025
机构 * Baichuan Inc.(北京百川科技有限公司) ; Peking University(北京大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments 14 pages,4 figures
机构 * Center for research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) ; Cornell University(康奈尔大学) ; Department of Electronics Engineering, Mehran University of Engineering & Technology(电子工程系,工程与技术大学) ; Department of Imaging Physics, The University of Texas MD Anderson Cancer Center(成像物理系,德克萨斯大学MD安德森癌症中心) ; Intelligent Transportation Systems, University of Tennessee(智能交通运输系统,田纳西大学) ; Department of Electrical Engineering, City University of Hong Kong(电气工程系,香港城市大学) ; Department of Mechanical Engineering, Bartin University(机械工程系,巴廷大学) ; Vector Institute, Toronto Canada(多伦多加拿大向量研究所) ; Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) ; Center for Data Science, New York University(数据科学中心,纽约大学) ; Meta Research(Meta 研究) ; Amazon Research(亚马逊研究) ; Centre for Artificial Intelligence Research and Optimization, Torrens University Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚) ; University Research and Innovation Center, Obuda University(研究与创新中心,奥布达大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * Institutes of Innovation for Future Society, Nagoya University, Japan(面向未来的创新研究所,名古屋大学,日本)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * MBZUAI ; The University of Sydney(悉尼大学) ; AI2Robotic ; Texas A&M University(德克萨斯大学) ; The University of Melbourne(墨尔本大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.CV
机构 * The Laboratory of Intelligent Collaborative Computing of UESTC(UESTC智能协同计算实验室) ; Ubiquitous Intelligence and Trusted Services Key Laboratory of Sichuan Province(四川省 Ubiquitous Intelligence and Trusted Services 重点实验室) ; Monash University(墨尔本大学) ; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Accepted to ICCV 2025