Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
机构 * Banking Academy of Vietnam(越南银行学院) ; Dublin City University(都柏林城市大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Banking Academy of Vietnam(越南银行学院) ; Dublin City University(都柏林城市大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * Waddle ; Seoul National University(首尔国立大学) ; Krafton ; UNIST(全南大学) ; SK Telecom(SK电信)
专题命中 视觉问答 :vision-language model(abstract);VLM(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted to EMNLP 2025 (Main Conference)
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Zhejiang Sci-Tech University(浙江科技学院)
专题命中 视觉问答 :visual question answering(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
Comments 23 pages
机构 * Cruise LLC (GM)(Cruise LLC(GM)) ; Northeastern University(东北大学) ; OpenAI(开放人工智能研究所) ; Meta ; University of Washington(华盛顿大学) ; Waymo LLC
专题命中 视觉推理 :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV、cs.LG
Comments Accepted by CoRL 2025
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) ; Department of Computer Science, Aalborg University(奥胡斯大学计算机科学系) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(title,abstract);分类 cs.AI
Comments 20 pages, 10 figures
机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI
Comments Conference on Robot Learning (CoRL) 2025. 50 pages and 30 figures. v2 is the camera-ready and includes a few more new experiments compared to v1
机构 * School of Future Technology, Shanghai University(未来技术学院,上海大学) ; School of Mechatronic Engineering and Automation, Shanghai University(机械电子工程与自动化学院,上海大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
Comments 14 pages,17 figures
机构 * Tencent Hunyuan Team(腾讯文言团队) ; Institute of Automation, CAS(中国科学院自动化研究所)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments 20 pages, 14 figures, 5 tables
机构 * Zhejiang University(浙江大学) ; Om AI Research(Om AI 研究所) ; Binjiang Institute of Zhejiang University(浙江大学滨江研究院)
专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/
机构 * School of Computer Science, Fudan University(复旦大学计算机学院) ; Weixin Group, Tencent(腾讯Weixin部门) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Key Lab of Intelligent Information Processing(上海智能信息处理重点实验室)
专题命中 视觉推理 :visual reasoning(abstract);multimodal large language model(abstract)
Comments Accepted to EMNLP 2025 Findings. The code and dataset are publicly available at https://github.com/hewei2001/ReachQA
机构 * University of Science and Technology of China(中国科学技术大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室)
专题命中 视觉推理 :vision language model(abstract);分类 cs.CV、cs.AI
Comments https://videorewardbench.github.io/
机构 * Toyota Technological Institute at Chicago (TTIC)(丰田技术研究所(芝加哥))
专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI
机构 * Duke University(杜克大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments The paper has been accepted to the 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), and selected for publication in the 2025 IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue
机构 * ONE Lab, HUST(华中科技大学 ONE 实验室) ; ONE Lab, HUST University of Maryland(华中科技大学 与 马里兰大学 ONE 实验室) ; University of Washington(华盛顿大学) ; Zhejiang University(浙江大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments Technical Report
机构 * Computing Science Department, Umeå University(乌梅大学计算科学系) ; CNRS IRL 2010 CROSSING(法国CNRS IRL 2010 CROSSING) ; IMT Atlantique(IMT阿蒂昂大学) ; National Institute of Informatics(日本信息机构)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学)
专题命中 视觉推理 :grounding(abstract)
Comments Accepted for IROS 2025
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG
Comments ICCV 2025 VLM 3d Workshop
机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) ; Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) ; Simula Research Laboratory, Norway(Simula研究实验室)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract)
机构 * Stanford University(斯坦福大学)
专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG
机构 * The University of Chicago(芝加哥大学) ; FabTrack
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments 20 pages, 5 figures
机构 * Intelligent Control \& Smart Energy (ICSE) Research Group, School of Engineering, University of Warwick, Coventry, CV4 7AL, UK
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Ho Chi Minh University of Science(胡志明市科学大学) ; Deakin University(德肯大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
Comments Published at ACMMM2025 (Dataset track)
机构 * University of Oxford(牛津大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Inserm U1331, Institut Curie Saint-Cloud, France(法国国家医学研究院U1331,圣克鲁医院) ; Aalto University Espoo, Finland(芬兰艾尔托大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) ; INESC-ID Lisboa(里斯本INESC-ID)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments 31 pages, 14 figures
机构 * Brown University(布朗大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Journal ref PVLDB, 18(11): 4073 - 4080, 2025
机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) ; NJUST(南京理工大学) ; College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) ; Tianjin University of Science and Technology(天津科技大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025
机构 * Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系) ; College of Literature, Science, and the Arts, University of Michigan(密歇根大学文学、科学与艺术学院)
专题命中 视觉定位与Grounding :VLM(abstract)
Comments 32 pages, 6 figures