What do vision-language models see in the context? Investigating multimodal in-context learning
机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL
机构 * Faculty of Computer Science and Engineering(计算机科学与工程学院) ; University Ss Cyril and Methodius(西里尔与美多西乌斯大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM
Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025
机构 * Department of Horticultural Sciences(园艺科学系) ; Gulf Coast Research and Education Center(墨西哥湾沿岸研究与教育中心) ; University of Florida(佛罗里达大学) ; Department of Agricultural and Biological Engineering(农业与生物工程系) ; Department of Soil, Water, and Ecosystem Sciences(土壤、水与生态系统科学系) ; Citrus Research and Education Center(柑橘研究与教育中心)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI
Comments 26 pages, 8 figures, and 2 tables
机构 * University of Chinese Academy of Science, Beijing,100190, China(中国科学院大学) ; Macquarie University(麦考瑞大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments Published in Pattern Recognition
机构 * Department of Computer Science(计算机科学系)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments 23 pages, 10 figures, 14 tables
机构 * Korea University(韩国大学) ; The Catholic University of Korea(韩国天主大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments Accepted to ICCV 2025. Code is available at: https://github.com/QuIIL/ICCV2025_Ano-NAViLa