arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4672 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2303.06571 2023-08-21 cs.CV 57%

Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language Models

Juncheng Li, Minghe Gao, Longhui Wei, Siliang Tang, Wenqiao Zhang, Mengze Li, Wei Ji, Qi Tian, Tat-Seng Chua, Yueting Zhuang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07648 2023-08-16 cs.CV 57%

Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval

Chaorui Deng, Qi Chen, Pengda Qin, Da Chen, Qi Wu

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments to be appeared in ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06628 2023-08-14 cs.CV cs.LG 57%

Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models

Zangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin, Xiangyu Yue, Yang You

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00353 2023-08-02 cs.CV 57%

Lowis3D: Language-Driven Open-World Instance-Level 3D Scene Understanding

Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, Xiaojuan Qi

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments submit to TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15983 2023-08-01 cs.MM 57%

Instance-Wise Adaptive Tuning and Caching for Vision-Language Models

Chunjin Yang, Fanman Meng, Shuai Chen, Mingyu Liu, Runtong Zhang

专题命中 图文多模态 :image-text(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00988 2023-07-24 eess.IV cs.CV cs.LG 57%

Continual Learning for Abdominal Multi-Organ and Tumor Segmentation

Yixiao Zhang, Xinyi Li, Huimiao Chen, Alan Yuille, Yaoyao Liu, Zongwei Zhou

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments MICCAI-2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10046 2023-07-20 cs.CV 57%

Divert More Attention to Vision-Language Object Tracking

Mingzhe Guo, Zhipeng Zhang, Liping Jing, Haibin Ling, Heng Fan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03339 2023-07-10 cs.CV 57%

Open-Vocabulary Object Detection via Scene Graph Discovery

Hengcan Shi, Munawar Hayat, Jianfei Cai

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17659 2023-07-03 cs.CV 57%

Zero-shot Nuclei Detection via Visual-Language Pre-trained Models

Yongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei, Maode Lai, Jianzhong Shou, Yubo Fan, Yan Xu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments This article has been accepted by MICCAI 2023,but has not been fully edited. Content may change prior to final publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16774 2023-06-30 cs.CL 57%

Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages

Yasmine Karoui, Rémi Lebret, Negar Foroutan, Karl Aberer

专题命中 图文多模态 :image-text(abstract);分类 cs.CL

Comments Accepted to ACL 2023 as short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14406 2023-06-30 cs.CV 57%

TCEIP: Text Condition Embedded Regression Network for Dental Implant Position Prediction

Xinquan Yang, Jinheng Xie, Xuguang Li, Xuechen Li, Xin Li, Linlin Shen, Yongqiang Deng

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16658 2023-06-30 cs.CV 57%

Prompt Ensemble Self-training for Open-Vocabulary Domain Adaptation

Jiaxing Huang, Jingyi Zhang, Han Qiu, Sheng Jin, Shijian Lu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15658 2023-06-28 cs.CV 57%

CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Xianhang Li, Zeyu Wang, Cihang Xie

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Tech Report. Code is available at https://github.com/UCSC-VLAA/CLIPA

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07279 2023-06-19 cs.CV 57%

Scalable 3D Captioning with Pretrained Models

Tiange Luo, Chris Rockwell, Honglak Lee, Justin Johnson

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Dataset link: https://huggingface.co/datasets/tiange/Cap3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14042 2023-06-16 cs.CV 57%

Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images

Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07831 2023-06-14 cs.CV 57%

Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images

Ming Y. Lu, Bowen Chen, Andrew Zhang, Drew F. K. Williamson, Richard J. Chen, Tong Ding, Long Phi Le, Yung-Sung Chuang, Faisal Mahmood

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03514 2023-06-12 cs.CV 57%

Recognize Anything: A Strong Image Tagging Model

Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, Lei Zhang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Homepage: https://recognize-anything.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02329 2023-06-12 cs.CV 57%

Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes

Alexandros Delitzas, Maria Parelli, Nikolas Hars, Georgios Vlassis, Sotirios Anagnostidis, Gregor Bachmann, Thomas Hofmann

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments The first two authors contributed equally. arXiv admin note: text overlap with arXiv:2304.06061

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03932 2023-06-08 cs.CV 57%

Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!

Zaid Khan, Vijay Kumar BG, Samuel Schulter, Xiang Yu, Yun Fu, Manmohan Chandraker

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09737 2023-06-08 cs.CV 57%

Position-guided Text Prompt for Vision-Language Pre-training

Alex Jinpeng Wang, Pan Zhou, Mike Zheng Shou, Shuicheng Yan

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments Camera-ready version, code is in https://github.com/sail-sg/ptp

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03678 2023-06-07 cs.CL 57%

On the Difference of BERT-style and CLIP-style Text Encoders

Zhihong Chen, Guiming Hardy Chen, Shizhe Diao, Xiang Wan, Benyou Wang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CL

Comments Natural Language Processing. 10 pages, 1 figure. Findings of ACL-2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16328 2023-05-29 cs.CL cs.LG 57%

Semantic Composition in Visually Grounded Language Models

Rohan Pandey

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CL

Comments Carnegie Mellon University Senior Thesis. arXiv admin note: substantial text overlap with arXiv:2212.10549

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15407 2023-05-25 cs.CV 57%

Balancing the Picture: Debiasing Vision-Language Datasets with Synthetic Contrast Sets

Brandon Smith, Miguel Farinha, Siobhan Mackenzie Hall, Hannah Rose Kirk, Aleksandar Shtedritski, Max Bain

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Github: https://github.com/oxai/debias-gensynth

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13782 2023-05-24 cs.CL 57%

Images in Language Space: Exploring the Suitability of Large Language Models for Vision & Language Tasks

Sherzod Hakimov, David Schlangen

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12777 2023-05-23 cs.CL 57%

Evaluating Pragmatic Abilities of Image Captioners on A3DS

Polina Tsvilodub, Michael Franke

专题命中 图文多模态 :image-text(abstract);分类 cs.CL

Comments 5 pages, 2 figures, to appear in the 61st Proceedings of the Association for Computational Linguistics (ACL 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10420 2023-05-18 cs.CV 57%

CLIP-GCD: Simple Language Guided Generalized Category Discovery

Rabah Ouldnoughi, Chia-Wen Kuo, Zsolt Kira

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09699 2023-05-18 cs.CV 57%

Mobile User Interface Element Detection Via Adaptively Prompt Tuning

Zhangxuan Gu, Zhuoer Xu, Haoxing Chen, Jun Lan, Changhua Meng, Weiqiang Wang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by CVPR23

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00788 2023-05-18 cs.CV 57%

Open-Vocabulary Point-Cloud Object Detection without 3D Annotation

Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, Shanghang Zhang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments I want to update this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06378 2023-05-18 cs.CV 57%

Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos

Teng Wang, Jinrui Zhang, Feng Zheng, Wenhao Jiang, Ran Cheng, Ping Luo

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06547 2023-05-04 eess.AS cs.SD 57%

Investigations in Audio Captioning: Addressing Vocabulary Imbalance and Evaluating Suitability of Language-Centric Performance Metrics

Sandeep Kothinti, Dimitra Emmanouilidou

专题命中 图文多模态 :cross-modal(abstract);分类 eess.AS

Comments Submitted to EUSIPCO 2023

详情

展开后加载摘要…

URL PDF HTML 收藏