Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
专题命中 视觉推理 :grounding(abstract)
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉推理 :grounding(abstract)
机构 * Institute for Infocomm Research (I 2 {}^{\text{2}} R), Agency for Science, Technology and Research (A*STAR)(信息与通信研究机构(I2R),科技研究局(A*STAR)) ; Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(前沿人工智能研究中心(CFAR),科技研究局(A*STAR))
专题命中 视觉推理 :grounding(abstract)
机构 * Southwest Jiaotong University(西南交通大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Tongji University(同济大学)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) ; Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) ; Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)
专题命中 视觉推理 :multimodal large language model(abstract)
Comments EMNLP 2025 Findings
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Sichuan University(四川大学) ; Shanghai Jiao Tong University(上海交通大学) ; University of Macau(澳门大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉推理 :multimodal large language model(abstract)
Comments 34 pages
机构 * Indian Institute of Technology Patna(印度理工学院帕纳巴分校) ; Banasthali Vidyapeeth University(班纳萨利大学) ; Pandit Deendayal Energy University(德英德能源大学) ; Manipal University Jaipur(马哈拉施特拉邦大学贾伊普尔分校) ; Dwarkadas J. Sanghvi College of Engineering(德瓦尔卡斯J.桑格维工程学院)
专题命中 视觉推理 :vision-language model(abstract)
Comments EMNLP MAINS 2025
机构 * Department of Robotics, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(机器人系,Mohamed bin Zayed人工智能大学)
专题命中 视觉推理 :visual reasoning(abstract)
Comments Submitted to IEEE for possible publication, under review
机构 * Allen Institute for AI(艾伦人工智能研究所) ; University of Washington(华盛顿大学)
专题命中 视觉推理 :grounding(abstract)
Comments Updated GR00T result to N1.5
专题命中 视觉推理 :grounding(abstract)
机构 * Zhejiang University(浙江大学)
专题命中 视觉推理 :grounding(abstract)
Comments Accepted for IROS 2025
专题命中 视觉推理 :grounding(abstract)
Comments 7 pages, 3 figures, 3 tables,
机构 * Qatar University(卡塔尔大学)
专题命中 视觉推理 :grounding(abstract)
机构 * Dongguk University(东国大学)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Yale University(耶鲁大学)
专题命中 视觉推理 :grounding(abstract)
Comments 12 pages, 4 figures, 2 tables. Extends our earlier framework on hierarchical narrative graphs with a semantic normalization module
机构 * Computational Linguistics, Department of Linguistics University of Potsdam(乌特雷赫特大学语言学系计算语言学部) ; German Research Center for Artificial Intelligence (DFKI), Berlin(德国人工智能研究中心(DFKI)柏林)
专题命中 视觉推理 :grounding(abstract)
Comments 17 pages
专题命中 视觉推理 :vision-language model(abstract)
专题命中 视觉推理 :multimodal large language model(abstract)
专题命中 视觉推理 :multimodal large language model(abstract)
机构 * Intelligent Autonomous Systems Group, Institute of Computational Visualisitics, University of Koblenz(智能自主系统组,计算可视化研究所,科隆大学)
专题命中 视觉推理 :vision-language model(abstract)
Comments Accepted and presented at the 1st German Robotics Conference (GRC); March 13-15, 2025, Nuremberg, Germany https://ras.papercept.net/conferences/conferences/GRC25/program/GRC25_ContentListWeb_3.html#sada_48
机构 * Stony Brook University(石溪大学)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨) ; Hospital Israelita Albert Einstein(埃因斯坦医院) ; Departamento de Pacientes Graves(重症患者部门) ; Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心) ; Departmento de Cirurgia(外科部门) ; Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) ; Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)
专题命中 视觉推理 :multimodal large language model(abstract)
专题命中 视觉推理 :multimodal large language model(abstract)
专题命中 视觉推理 :vision-language model(abstract)
Comments Code Available: https://github.com/chan0park/SelfReVision
机构 * Beihang University(北航大学) ; Guangxi Normal University(广西师范大学) ; University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 视觉推理 :grounding(abstract)
机构 * Huawei Technologies(华为技术有限公司) ; HKUST(香港科技大学)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Hong Kong University of Science and Technology(香港科技大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 视觉推理 :vision-language model(abstract)
Comments Project Page: https://vlm2-bench.github.io/ Camera Ready version
专题命中 视觉推理 :vision-language model(abstract)
Comments VIS 2025
机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; SCUT ; SEU ; BAAI(北京人工智能研究院)
专题命中 视觉推理 :vision-language model(abstract)
机构 * Key Laboratory of Embedded System and Service Computing, Ministry of Education, Tongji University(嵌入式系统与服务计算重点实验室,教育部,同济大学) ; School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)
专题命中 视觉推理 :grounding(abstract)
Comments 20 pages; Accepted to ACL 2025 Main
机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通讯学院,华中科技大学) ; School of Basic Medicine, Huazhong University of Science and Technology(基础医学院,华中科技大学)
专题命中 视觉推理 :multimodal large language model(abstract)
Comments Multimodal Benchmark, Project Url: https://github.com/micdz/MANBench, ACL2025 Findings