[CLS] Token is All You Need for Zero-Shot Semantic Segmentation
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 8 pages,6 figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 8 pages,6 figures
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 11 Pages, 6 figures, 8 tables, Accepted in Earth Vision (CVPR 2023)
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.CV
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.CV
Comments Accepted to CVPR 2023
专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV
Comments accepted by CVPR23
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments Accepted to ICLR 2023
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments Accepted by AAAI2023
专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments Technical Report
专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV
Comments Accepted by AAAI2023
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 11 pages (including Supplementary Materials); Accepted to ACM MM 2022
Journal ref ACM International Conference on Multimedia. 2022. 3587-3597
专题命中 VLM训练与架构 :visual question answering(abstract);分类 cs.CV
Comments This paper is accepted for publication as a REGULAR paper in the IEEE Transactions on Multimedia
专题命中 VLM训练与架构 :MLLM(abstract);分类 cs.AI
Comments Our code, models, and data (e.g., integration corpus and extended datasets) are available: https://github.com/yifan-h/Multilingual_Space
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments To appear at NeurIPs 2022, Camera Ready with Typos fixed
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.CV
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 20 pages, 23 figures
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments BMVC 2022
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments Accepted by NeurIPS2022. Code is available at https://github.com/xk-huang/OrdinalCLIP
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.AI
Comments Accepted to NAACL 2021. The Spider-Realistic dataset is available at https://doi.org/10.5281/zenodo.5205322
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 4 pages, 4 figures, under review
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.LG
Journal ref Advances in Neural Information Processing Systems 2020
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments Accepted by ACL 2022 (long paper)
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV
Comments 14 pages
专题命中 VLM训练与架构 :visual question answering(abstract);分类 cs.CV
Comments ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)
专题命中 VLM训练与架构 :visual question answering(abstract);分类 cs.LG
Comments NeurIPS 2021; Our code is publicly available at https://github.com/szhang42/alignment_attention
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.AI
Comments EACL 2021
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.AI
Comments Preprint version; final version available at http://ieeexplore.ieee.org/ IEEE Transactions on Cognitive and Developmental Systems (Accepted) DOI: 10.1109/TCDS.2017.2754143
Journal ref IEEE Transactions on Cognitive and Developmental Systems 10 (4), 1005-1022, 2018
专题命中 VLM训练与架构 :grounding(abstract);分类 cs.CV
Comments ECCV2020, 18 pages, 6 figures
专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV
Comments accepted by AAAI-2020. arXiv admin note: text overlap with arXiv:1909.11740 by other authors