Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
机构 * Yonsei University(延世大学)
专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV
Comments 17 pages
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Yonsei University(延世大学)
专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV
Comments 17 pages
机构 * Xi'an Jiaotong University(西安交通大学) ; Queen Mary University of London(伦敦女王玛丽大学) ; Center for Multimodal AI(多模态人工智能中心) ; Digital Environment Research Institute(数字环境研究院)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments accepted for publication at ACM Multimedia (ACM MM) 2025
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
机构 * UC Davis(加州大学戴维斯分校) ; University of South Florida(佛罗里达州立大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
机构 * Beijing Jiaotong University(北京交通大学) ; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
专题命中 视觉定位与Grounding :MLLM(title);分类 cs.CV
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; School of Computer Science, Peking University(北京大学计算机学院) ; School of EECS, Peking University(北京大学电子工程学院)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments ICCV 2025
机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(人工智能与机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments Accepted to IROS 2025. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align
机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) ; Linköping University(林奈大学) ; Australian National University(澳大利亚国立大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments 20 pages, 13 figures
机构 * Artificial Intelligence Creative Research Lab, ETRI University of Science(人工智能创意研究实验室,ETRI大学) ; Field Robotics Research Section, ETRI University of Science(机器人领域研究部,ETRI大学) ; Artificial Intelligence Creative Research Lab, ETRI Daejeon, South Korea(人工智能创意研究实验室,ETRI大田,韩国)
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV
Comments Accepted at the 17th IEEE International Conference on Advanced Computational Intelligence (ICACI 2025)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments Accepted by IJCAI 2025 Survey Track
专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG
Comments 13 pages, 12 figures
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments Accepted at ICLR 2025, code: https://github.com/aluo-x/BrainSAIL
机构 * Division of Medical Image Computing, German Cancer Research Center, Heidelberg, Germany(德国癌症研究中心医学图像计算部) ; Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany(海德堡大学数学与计算机科学学院) ; Medical Faculty Heidelberg, University of Heidelberg, Heidelberg, Germany(海德堡大学医学学院)
专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV
Comments Accepted at ECCV 2024 Workshop on Emergent Visual Abilities and Limits of Foundation Models & Medical Imaging with Deep Learning 2025
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments 14 pages, 2 figures
机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人类-机器混合增强智能重点实验室) ; National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) ; Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) ; Xi’an Jiaotong University(西安交通大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
机构 * Institute Science of Tokyo(东京科学研究所)
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV
机构 * Mila-Québec, Université de Montréal(蒙特利尔大学魁北克分校) ; Carnegie Mellon University(卡内基梅隆大学) ; Deutsches Forschungszentrum für künstliche Intelligenz (DFKI)(德国人工智能研究中心) ; University of Pennsylvania(宾夕法尼亚大学) ; Next Generation Analytics and Modulo Bio(下一代分析与Modulo Bio) ; ServiceNow Research(ServiceNow研究) ; Mila-Québec, Université Laval(魁北克蒙特利尔大学拉瓦尔分校)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG
Comments 39 pages, 8 figures; CLeaR 2025
机构 * Sun Yat-sen University, Guangzhou, China(中山大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV
Comments Accepted at ICIP2025 Dataset and Benchmark Track
机构 * Nanjing University(南京大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Accepted by ICML 2025, 21 pages
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Presented at the 15th International Conference on Computational Creativity (ICCC'24)
Journal ref Proceedings of the Fifteenth International Conference on Computational Creativity (2024) 101-106
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Oral at NAACL 2025 Main conference. Albuquerque, USA. Apr 29 - May 4, 2025. 19 pages, 9 figures, 7 tables
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV
Comments 8 figures; 6 tables
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Add repair model ablation, update related work
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments Updated in Jan. 2025, In Proceedings of the European Conference on Computer Vision 2022 [ECCV 2022], 27 pages
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Accepted at COLING 2025