Towards Understanding Camera Motions in Any Video
专题命中 视觉问答 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Project site: https://linzhiqiu.github.io/papers/camerabench/
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Project site: https://linzhiqiu.github.io/papers/camerabench/
机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV
Comments Accepted to ICCV Workshop 2025
机构 * Department of Artificial Intelligence, Xiamen University(人工智能学院,厦门大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
Comments Accepted by ACM MM'25
机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) ; Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) ; Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 6 figures
机构 * Google DeepMind(谷歌DeepMind)
专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract)
Comments To be published in Proceedings of the 9th Conference on Robot Learning (CoRL). 34 pages, 10 figures
机构 * University of Colorado Colorado Springs(科罗拉多州立大学) ; University of Tennessee–Oak Ridge Innovation Institute(田纳西大学-橡树岭创新研究所) ; Xi’an Jiaotong University(西安交通大学) ; University of West Bohemia(西波维亚大学) ; Honda Research Institute USA(本田美国研究机构) ; University of Notre Dame(诺特丹大学) ; Can Tho University(庆和大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments 11 pages, 2 figures, Accepted to ICCV 2025 Workshop on Out-of-Label Hazards in Autonomous Driving (2COOOL)
机构 * Meituan(美团)
专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV
Comments 24 pages
机构 * College of Information Science and Engineering, Northeastern University, Shenyang, China(信息科学与工程学院,东北大学,沈阳,中国) ; School of Mechanical Engineering and Automation, Northeastern University, Shenyang, China(机械工程与自动化学院,东北大学,沈阳,中国) ; Enterprise AI, Midea AI Innovation Center, Foshan, China(企业人工智能,美的人工智能创新中心,佛山,中国) ; Foshan Graduate School of innovation, Northeastern University, Foshan, China(佛山创新研究生学院,东北大学,佛山,中国)
专题命中 GUI与屏幕智能体 :grounding(abstract);分类 cs.CV
机构 * Institute for Systems and Robotics(系统与机器人研究所) ; University of Lisbon(里斯本大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);InternVL(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) ; School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) ; Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Renmin University of China(中国人民大学) ; National Yang Ming Chiao Tung University ; Taipei Veterans General Hospital(台北荣民总医院) ; School of Biomedical Engineering, Shenzhen University(深圳大学生物医学工程学院)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 19 pages, 5 figures, 3 tables
机构 * ISTBI Fudan University(ISTBI 复旦大学) ; Huashan Hospital Fudan University(复旦大学华山医院) ; BME & CBIS Rensselaer Polytechnic Institute(生物医学工程与生物信息学研究中心罗切斯特理工学院)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments 14 pages, 9 figures
Journal ref Visual Computing for Industry, Biomedicine, and Art, 7, 20, 2024
专题命中 其他VLM :MLLM(abstract)
Comments 8 pages, 5 figures