K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
机构 * Baichuan Inc(百川公司) ; Peking University(北京大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Baichuan Inc(百川公司) ; Peking University(北京大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) ; Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(化学与生物工程系,香港科学与技术大学) ; Division of Life Science, The Hong Kong University of Science and Technology(生命科学系,香港科学与技术大学)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments For data and code, see: https://huggingface.co/datasets/slyipae1/MedBookVQA and https://github.com/slyipae1/MedBookVQA
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments More visualizations on Homepage: https://alpha-innovator.github.io/OmniCaptioner-project-page and Official code: https://github.com/Alpha-Innovator/OmniCaptioner
机构 * Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理关键实验室) ; School of Computer Science, Fudan University(复旦大学计算机科学学院)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * HKU(香港大学) ; SCUT(华南理工大学) ; SJTU(上海交通大学) ; PKU(北京大学) ; Allen AI(AllenAI)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments CVPR 2025 Camera Ready Version. Project page: https://vl-rewardbench.github.io
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; AD Lab, CaiNiao Inc., Alibaba Group(阿里集团 Cainiao 事业部实验室)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Project Page: https://PixelThink.github.io
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Xidian University(西安电子科技大学) ; Sun Yat-sen University(中山大学) ; The University of Sydney(悉尼大学) ; INSAIT, Sofia University(INSAIT,索菲亚大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Computer Science and Engineering(计算机科学与工程) ; Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) ; UC San Diego(圣地亚哥大学) ; Harvard SEAS(哈佛大学工程与应用科学学院) ; MIT Mathematics(MIT数学系) ; Halıcıoğlu Data Science Institute(Halıcıoğlu数据科学研究所) ; Harvard CMSA(哈佛大学计算机科学与应用数学系)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology, China(计算机科学与工程学院,南京理工大学)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI
机构 * Tencent YouTu Lab(腾讯YouTu实验室)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
专题命中 视觉推理 :grounding(abstract);分类 cs.CV
Comments Project page: https://aim-uofa.github.io/OmniR1
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Institute of Automation, CAS(中国科学院自动化研究所)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * Hong Kong Polytechnic University(香港理工大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Microsoft(微软公司)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * R&D Lab F-Initiatives(F-Initiatives研发实验室) ; Université d’Orleans(奥尔良大学) ; Université Sorbonne Paris Nord(巴黎-索邦大学) ; University of Dubai(迪拜大学) ; IULM University(IULM大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Under review
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments CVPR 2025. Project website: https://yufu-wang.github.io/phmr-page
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Zhejiang University(浙江大学) ; School of Science and Engineering, The Chinese University of Hong Kong(香港中文大学科学与工程学院) ; Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; Network Science Institute, Northeastern University(东北大学网络科学研究所)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments 5 pages, 2 figures; accepted by IEEE VIS 2024 (https://ieeevis.org/year/2024/program/paper_v-short-1177.html)
Journal ref 2024 IEEE Visualization and Visual Analytics (VIS)
机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) ; Nanyang Technological University(南洋理工大学) ; WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学) ; Shanghai Jiao Tong University(上海交通大学) ; Nanjing University(南京大学) ; Fudan University(复旦大学) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉推理 :vision language model(abstract);分类 cs.CV
Comments ACL 2025 Findings
专题命中 视觉推理 :VLM(abstract);分类 cs.CV
Comments Best Student Paper Award at IEEE International Conference on Computer Supported Cooperative Work in Design, 2025
机构 * Kai Zhang 1(某机构) ; Xingyu Chen 3(某机构) ; Xiaofeng Zhang 2(某机构)
专题命中 视觉推理 :LLaVA(abstract);分类 cs.CV
机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, and Institute of Artificial Intelligence, Wuhan University, China(计算机学院、多媒体软件国家工程研究中心及人工智能研究所、武汉大学) ; College of Computing & Data Science at Nanyang Technological University(南洋理工大学计算与数据科学学院)
专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV
机构 * Shanghai University(上海大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉推理 :VLM(abstract);分类 cs.AI
机构 * Embodied AI and Robotics (AIR) Lab New York University Abu Dhabi(embodied AI 和机器人(AIR)实验室 新 York 大学阿布扎赫德)
专题命中 视觉推理 :MLLM(abstract);分类 cs.CV
Comments Project website and code: https://dktpt44.github.io/LV-GPT/
机构 * School of Cyber Science and Engineering, Xi’an Jiaotong University(网络安全与工程学院,西安交通大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI
Comments 11 pages, 5 figures, 3 tables, submitted to IEEE OJCS
机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Alibaba Group(阿里巴巴集团) ; DAMO Academy, Alibaba Group(阿里巴巴集团大模型学院) ; Nanjing Institute of Software Technology(南京软件技术研究所) ; Nanjing University of Posts and Telecommunications(南京邮电大学) ; Hohai University(河海大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments Code: https://github.com/zwq2018/embodied_reasoner Dataset: https://huggingface.co/datasets/zwq2018/embodied_reasoner
机构 * University of Washington(华盛顿大学) ; Allen Institute for AI(人工智能研究院)
专题命中 视觉推理 :LLaVA(abstract);分类 cs.CV
Comments CVPR 2025