Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
机构 * Meta ; KU Leuven(鲁汶大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments ICCV 2025 CDEL Workshop
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Meta ; KU Leuven(鲁汶大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments ICCV 2025 CDEL Workshop
机构 * Faculty of Electrical and Electronics Engineering, University of Engineering(电气电子工程学院,工程大学) ; Department of Engineering Sciences, University of Agder(工程科学系,阿格德大学) ; Sheikh Zayed Institute for Pediatric Surgical Innovation, Children’s National Hospital(谢赫扎耶德小儿外科创新研究所,儿童医院) ; School of Medicine and Health Sciences, George Washington University(医学与健康科学学院,乔治华盛顿大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.LG
机构 * organization= Agricultural \& Biological Engineering, Purdue University , country= USA ; organization= Environmental \& Ecological Engineering, Purdue University , country= USA ; organization= Davidson School of Chemical Engineering, Purdue University , country= USA ; organization= Mechanical \& Mechatronic Engineering, University of Technology Sydney (UTS) , country= Australia
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025
机构 * University of Southern California(南加州大学) ; National Technical University of Athens(雅典国立技术大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments 16 pages, 12 figures, 3 tables
Journal ref Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5513-5528, Albuquerque, New Mexico, April 2025
机构 * Beijing Normal University(北京师范大学) ; University of Chinese Academy of Sciences(中国科学院大学) ; Fudan University(复旦大学) ; Central University of Finance and Economics(中央财经大学) ; Tsinghua University(清华大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments ICCV2025
机构 * Department of Computer Science and Engineering, Sungkyunkwan University(全南大学计算机科学与工程系) ; Acryl Inc.(阿克罗尔公司)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments 21 pages. NeurIPS 2024
Journal ref Advances in Neural Information Processing Systems 37, 67034-67060, 2024
机构 * University of Applied Sciences and Arts of Southern Switzerland(应用科学与艺术大学(南瑞士)) ; Dalle Molle Institute for Artificial Intelligence(达勒莫勒人工智能研究所) ; University of Florence(佛罗伦萨大学) ; Department of Mathematics and Computer Science Ulisse Dini(数学与计算机科学系(乌利塞·迪尼))
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,阿基米德大学塞萨洛尼基分校) ; Information Technology Institute, Centre for Research & Technology Hellas(信息科技研究所,希腊研究中心)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments 58 page
机构 * Czech Institute of Informatics, Robotics and Cybernetics at the Czech Technical University in Prague(捷克信息技术、机器人与自动化研究所(捷克技术大学)) ; Inria, École normale supérieure, CNRS, PSL Research University(法国国家信息与自动化研究所、法国高等师范学校、法国国家科学研究中心、巴黎-萨克勒大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Accepted at ICCV 2025. Erratum: An earlier version reported ablations (Table 6 & Fig. 6) with pre-training on a 50k subset of HowToGround1M + fine-tuning on iGround. In the ICCV camera-ready, Table 6 already used the full dataset, but Fig. 6 and a sentence in the text were mistakenly left on 50k. All now use the full HowToGround1M
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV
机构 * Mercedes-Benz AG, Germany(梅赛德斯-奔驰集团,德国) ; Technical University of Munich, Germany(慕尼黑技术大学,德国) ; Ferdinand-Steinbeis-Institut der Steinbeis-Stiftung, Germany(施坦贝格基金会费尔迪南-斯坦贝格研究所,德国)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI
机构 * University of Pennsylvania(宾夕法尼亚大学) ; Archimedes, Athena RC(阿基米德、阿提卡RC)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
Comments ICCV 2025
机构 * PES University(PES大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Media Arts & Technology UC Santa Barbara(媒体艺术与技术大学圣芭芭拉分校)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments to be published in IEEE VISAP 2025
机构 * University of Glasgow(格拉斯哥大学) ; Jadavpur University(贾瓦德普尔大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments I want to revisit some of the experiments in this paper, specifically figure 5
机构 * University of Oxford(牛津大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Inserm U1331, Institut Curie Saint-Cloud, France(法国国家医学研究院U1331,圣克鲁医院) ; Aalto University Espoo, Finland(芬兰艾尔托大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) ; INESC-ID Lisboa(里斯本INESC-ID)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments 31 pages, 14 figures
机构 * Brown University(布朗大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Journal ref PVLDB, 18(11): 4073 - 4080, 2025
机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) ; NJUST(南京理工大学) ; College of Mathematics, Taiyuan University of Technology(数学学院,太原科技大学) ; Tianjin University of Science and Technology(天津科技大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments Extension of our Findings of EMNLP 2023 & ACL 2024 paper, IEEE Transactions on Multimedia accepted on July 19, 2025
机构 * University of Colorado Colorado Springs(科罗拉多州立大学) ; University of Tennessee–Oak Ridge Innovation Institute(田纳西大学-橡树岭创新研究所) ; Xi’an Jiaotong University(西安交通大学) ; University of West Bohemia(西波维亚大学) ; Honda Research Institute USA(本田美国研究机构) ; University of Notre Dame(诺特丹大学) ; Can Tho University(庆和大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments 11 pages, 2 figures, Accepted to ICCV 2025 Workshop on Out-of-Label Hazards in Autonomous Driving (2COOOL)
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Seoul National University(首尔国立大学) ; NVIDIA(NVIDIA公司)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/
机构 * Department of Computer Science, University of Virginia(大学计算机科学系) ; University of Virginia School of Medicine(弗吉尼亚大学医学院) ; Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) ; Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) ; Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
专题命中 视觉定位与Grounding :vision language model(abstract);分类 cs.CV
Comments Accepted at BMVC 2025
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
Comments 12 pages, 7 figures, 5 tables
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI